Independent DIB implementation resource — not affiliated with or endorsed by the U.S. Department of WarView the official DoW campaign ↗
OT-08OPERATIONAL TECHNOLOGYOFFICIAL TITLEREVIEWED — PENDING SME SIGN-OFF

System Resiliency

Resiliency is the ability to keep operating, or return quickly, when something fails — whether a disk, a controller, or an attack. In OT that means backups of controller logic and configurations, spares and redundancy for critical components, defined safe-state and manual-operation fallbacks, and recovery that has actually been tested against how long production can be down.

Independent interpretation

The title above is the official campaign practice name. Everything else on this page — the sequencing, the actions, the maturity ladder, the validation checks, the evidence guidance, and the framework mappings — is independent analysis by the Brilliant at the Basics Resource Center. It carries no official status and is not endorsed by the U.S. Department of War. The official campaign ↗ remains authoritative.

EXPLAINER · 5 SCENES · ≈40 SEC · CAPTIONS, NO AUDIO

OT-08 in 40 seconds

The problem, the plain-words meaning, three key moves, and what “done” looks like.

Official intent

What the campaign asks for

Engineer system resiliency so you keep producing — or recover fast — when OT systems fail. The official source remains authoritative.

Read the official campaign ↗

Resiliency is the ability to keep operating, or return quickly, when something fails — whether a disk, a controller, or an attack. In OT that means backups of controller logic and configurations, spares and redundancy for critical components, defined safe-state and manual-operation fallbacks, and recovery that has actually been tested against how long production can be down.

Who this applies to: Every environment where downtime has a cost. Resiliency work is usually justified by availability long before it is justified by security.

Why it matters

OT failures have physical and financial consequences measured in downtime, scrap, and sometimes safety. Resiliency limits the impact of the failure you could not prevent — the difference between a brief switch to a backup and a multi-day outage while you rebuild a controller from memory. Tested recovery is what turns a plan into a capability.

Risks this reduces

  • Extended downtime rebuilding a controller whose logic was never backed up
  • Single points of failure with no spare and a long lead time
  • Redundancy that has never been failed over and does not engage cleanly
  • Recovery expectations that were never checked against what the business can survive
Coordinate before touching production

Testing failover, recovery, or safe-state transitions on live systems can itself cause a disruption. Exercise recovery in a lab or during planned windows with the process owner, validate that safe-state and manual fallbacks behave as expected, and never assume an untested redundancy will engage cleanly under real failure.

Who owns it

Primary owner

Plant / OT leader

Supporting

OT engineer, Maintenance, Vendors

Effort

High

Cost band

High

Ownership is a named person, not a department. If nobody can be named, that is the first finding.

Dependencies and prerequisites

Leans on: OT-02Validated OT asset inventory

  • A validated inventory that identifies critical assets and their process dependencies
  • Operator agreement on what safe state and manual operation actually look like
  • Somewhere off the device to store logic, configuration, and set-point backups

Action timeline

First 24 hoursConfirmation and discovery. Nothing here needs procurement.
  • Confirm backups exist for controller logic and configurations
  • List critical single points of failure
By day 14The first changes that measurably reduce exposure.
  • Confirm that controller logic, configuration, and set points are backed up for your most critical line, and store a copy off the device
  • Write down the safe-state and manual-operation fallback with the operators who would use it
By day 30Coverage across the intended scope.
  • Define acceptable downtime for critical processes
  • Document safe-state and manual-operation fallbacks
By day 90Operating, measured, and reviewable.
  • Address the top single points of failure with spares or redundancy
  • Test-restore a controller configuration
  • Run a failure/recovery exercise against the downtime target
  • Feed gaps into a resiliency investment plan

The 90-day target for this practice is the Measured level below: coverage and effectiveness are reported, and exceptions are handled rather than accumulated.

Step-by-step implementation

  1. Back up controller logic, configurations, and set points, and store copies safely off the device.
  2. Identify critical single points of failure and the processes that cannot tolerate downtime.
  3. Define acceptable downtime and document safe-state and manual-operation fallbacks with operators.
  4. Add redundancy or spares for the most critical components.
  5. Test recovery against the downtime target and exercise realistic failure and attack scenarios.

What good looks like

Seven levels, used identically across every practice, scorecard, and download on this site. The distinction that matters most is between having a tool, deploying it to the correct scope, and operating it consistently.

  1. Absent

    No controller backups; recovery would rely on someone's memory or a vendor's availability.

    Is there anything at all — a tool, a document, a person who owns it?
  2. Documented

    Critical single points of failure are identified and acceptable downtime is defined with the business.

    Is the intent written down, with a named owner and a scope?
  3. Configured

    Logic and configuration backups are being taken and stored away from the devices they came from.

    Is it switched on and set up somewhere — even if only in part of the estate?
  4. Deployed

    Backups cover critical control systems, safe-state and manual fallbacks are documented, and spares exist for the top failure points.

    Does it cover everything in scope, with the exceptions written down?
  5. Operating

    Backups are refreshed after every change, and operators know and can execute the fallbacks.

    Does it keep working through a normal month without manual rescue?
  6. Measured

    Recovery is tested against the downtime target and the gap between tested and required recovery is reported.

    Can you state a number for coverage or effectiveness, and show the trend?
  7. Governed

    An owner runs scenario exercises, reviews objectives against plant change, and gaps drive funded investment.

    Is there an accountable owner, a review cadence, and retained evidence?

Validation procedures

Until these pass, the practice is configured — not deployed.

  • Test-restore a controller's logic/configuration and confirm the process resumes correctly.
  • Verify a defined safe-state or manual fallback exists and is understood by operators.
  • Confirm the last recovery exercise met — or exposed a gap against — the downtime target.

Evidence to retain

Governance

OT resiliency/contingency plan with downtime targets and fallbacks

Configuration

Controller backup inventory and redundancy/spares list

Operations

Recovery-exercise reports and single-point-of-failure remediation

Validation

Test-restore results measured against the downtime target

Retaining these supports your own assurance and gives a reviewer something concrete to examine. It does not constitute an assessment or satisfy a contractual requirement on its own.

Operating metrics

MetricHow it is calculatedDirectional target
Controller backup coverageCritical controllers with a current logic and configuration backup ÷ critical controllers.100%, refreshed after every change.
Tested recovery timeElapsed time of the most recent restore test, compared with the agreed acceptable downtime.Within the target; shortfalls documented and funded.
Unmitigated single points of failureCritical components with neither redundancy nor an obtainable spare.Zero, or an accepted risk with a named owner and a date.

Targets are directional guidance for your own programme, not compliance thresholds.

Common failure modes

What looks done but is not

Backing up servers but not PLC logic and set points, redundancy that has never been failed-over to test it, and a recovery plan that assumes parts and expertise you cannot get quickly. Untested redundancy is a hope with a wiring diagram.

Two sized paths

Small businessLittle or no dedicated IT staff

Back up the logic and configuration for every controller you would struggle to rebuild, keep a copy somewhere that is not the plant floor, and restore one of them to prove it works. Then write down which spare parts have long lead times. That is most of the practical value.

Mature environmentDedicated security capability

Exercise recovery end-to-end for realistic failure and attack scenarios, measure against per-process downtime objectives, verify redundancy by actually failing over during planned windows, and let measured gaps drive the resiliency investment plan.

Tool categories

Categories, not recommendations. This site ranks no vendors and accepts no paid placement.

Controller logic and configuration backup toolingOff-device, offline backup storageRedundant hardware and critical spares programmeRecovery runbooks and safe-state procedures
Change control

Backups age out the moment logic changes, so make a backup refresh part of the change record itself. Test failover and recovery in a lab or planned window with the process owner — an untested redundancy that fails during a real event is worse than a known gap.

Framework mappings

Independent mappings are aids to your own analysis — not authoritative equivalence, not coverage, and not a compliance determination. Read the caveat on every row before using it.

NIST SP 800-82 Rev. 3§3.3.9 — Develop a Recovery and Restoration Capability
DirectHigh confidence

Why: The publication addresses recovery and restoration for control systems including configuration backups and degraded-mode operation.

What this does not claim: 800-82 is guidance, not a compliance standard.

NIST SP 1339 (June 2026)OT Backup Quick Start Guide
DirectHigh confidence

Why: The guide is specifically about regular OT backups, their integration with change management, testing, and recovery exercises — the operational core of this practice.

What this does not claim: A quick-start guide, advisory by nature. It informs how backups and recovery testing are run; it imposes no requirement and does not cover redundancy, spares, or safe-state design.

NIST CSF 2.0RC.RP-01 / PR.DS-11
DirectModerate confidence

Why: Recovery-plan execution and backup creation are the two outcomes this practice produces.

What this does not claim: The CSF describes outcomes, not testable controls.

IEC 62443-3-3:2013SR 7.3 / SR 7.4
DirectLow confidence

Why: Control system backup and recovery/reconstitution are explicit system requirements in the availability family.

What this does not claim: IEC 62443 is a voluntary industrial standard; conformance is a separate, scoped exercise. The specific requirement identifiers are pending verification against the licensed normative text and should be treated as directional until that check completes.

Method: Read against the primary source text, then classified by relationship type and confidence. No automated mapping tool was used. How mappings are made →

NIST SP 800-171 relationships have their own requirement-level section below, with links into the full mapping experience.

Related NIST SP 800-171 requirements

Requirement-level relationships from the site’s independent 800-171 mapping. Each one names what it does and does not claim — a practice supports implementation of a requirement; it never satisfies one by itself. Expand a row for the rationale and caveat.

NIST SP 800-171 Rev. 2

3.8.9Backup CUI confidentialityPartial implementation supportLow

Why: The practice's off-device stores of controller logic, configuration, and set points are backup media, and where historian data, engineering workstation images, or project files in those stores contain CUI, protecting the copies advances this requirement.

What this does not claim: Directional only: most controller backups contain no CUI, so this mapping rarely applies. The practice's aim is availability — recovering production fast — and confidentiality protection of the backup store is not its default posture; where CUI is present, encryption and access limitation must be added deliberately, and adding them must not leave the plant unable to reach its own recovery files during an outage.

Review status: Technical review complete.

Open the 3.8.9 page →

NIST SP 800-171 Rev. 3

03.08.09System Backup — Cryptographic ProtectionPartial implementation supportModerate

Why: System resiliency for OT leans on maintained configuration and system backups for critical production assets, and those backup stores fall inside this requirement's scope the moment they contain CUI — protecting them is part of the practice's recovery posture.

What this does not claim: The practice's center of gravity is availability — keeping production running and restoring it fast — while this requirement is about the confidentiality of what the backups contain. Encrypting OT configuration backups and restricting access to the repository are additional steps the resiliency work does not automatically include, and older OT backup tooling may not support encryption at all, forcing compensating protection of the storage location itself.

Review status: Pending NIST SME review.

Open the 03.08.09 page →

Review status

Technical reviewReviewed — pending SME sign-off
Editorial reviewReviewed
Reviewed byinDirectIT practitioner review — CUI security and NIST SP 800-171 engineering
Last reviewed
Official source verified
Content version1.0

This guide has been reviewed by practitioners but is awaiting sign-off from a subject-matter expert in this specific domain. Treat the safety and change-control guidance as a floor, not a ceiling, and validate it against your own process and vendor requirements.