Official intent
Engineer system resiliency so you keep producing — or recover fast — when OT systems fail. The official source remains authoritative.
Read the official campaign ↗Resiliency is the ability to keep operating, or return quickly, when something fails — whether a disk, a controller, or an attack. In OT that means backups of controller logic and configurations, spares and redundancy for critical components, defined safe-state and manual-operation fallbacks, and recovery that has actually been tested against how long production can be down.
Who this applies to: Every environment where downtime has a cost. Resiliency work is usually justified by availability long before it is justified by security.
Why it matters
OT failures have physical and financial consequences measured in downtime, scrap, and sometimes safety. Resiliency limits the impact of the failure you could not prevent — the difference between a brief switch to a backup and a multi-day outage while you rebuild a controller from memory. Tested recovery is what turns a plan into a capability.
Risks this reduces
- Extended downtime rebuilding a controller whose logic was never backed up
- Single points of failure with no spare and a long lead time
- Redundancy that has never been failed over and does not engage cleanly
- Recovery expectations that were never checked against what the business can survive
Testing failover, recovery, or safe-state transitions on live systems can itself cause a disruption. Exercise recovery in a lab or during planned windows with the process owner, validate that safe-state and manual fallbacks behave as expected, and never assume an untested redundancy will engage cleanly under real failure.
Who owns it
Plant / OT leader
OT engineer, Maintenance, Vendors
High
High
Ownership is a named person, not a department. If nobody can be named, that is the first finding.
Dependencies and prerequisites
Leans on: OT-02 — Validated OT asset inventory
- A validated inventory that identifies critical assets and their process dependencies
- Operator agreement on what safe state and manual operation actually look like
- Somewhere off the device to store logic, configuration, and set-point backups
Action timeline
- Confirm backups exist for controller logic and configurations
- List critical single points of failure
- Confirm that controller logic, configuration, and set points are backed up for your most critical line, and store a copy off the device
- Write down the safe-state and manual-operation fallback with the operators who would use it
- Define acceptable downtime for critical processes
- Document safe-state and manual-operation fallbacks
- Address the top single points of failure with spares or redundancy
- Test-restore a controller configuration
- Run a failure/recovery exercise against the downtime target
- Feed gaps into a resiliency investment plan
The 90-day target for this practice is the Measured level below: coverage and effectiveness are reported, and exceptions are handled rather than accumulated.
Step-by-step implementation
- Back up controller logic, configurations, and set points, and store copies safely off the device.
- Identify critical single points of failure and the processes that cannot tolerate downtime.
- Define acceptable downtime and document safe-state and manual-operation fallbacks with operators.
- Add redundancy or spares for the most critical components.
- Test recovery against the downtime target and exercise realistic failure and attack scenarios.
What good looks like
Seven levels, used identically across every practice, scorecard, and download on this site. The distinction that matters most is between having a tool, deploying it to the correct scope, and operating it consistently.
- Absent
No controller backups; recovery would rely on someone's memory or a vendor's availability.
Is there anything at all — a tool, a document, a person who owns it? - Documented
Critical single points of failure are identified and acceptable downtime is defined with the business.
Is the intent written down, with a named owner and a scope? - Configured
Logic and configuration backups are being taken and stored away from the devices they came from.
Is it switched on and set up somewhere — even if only in part of the estate? - Deployed
Backups cover critical control systems, safe-state and manual fallbacks are documented, and spares exist for the top failure points.
Does it cover everything in scope, with the exceptions written down? - Measured
Recovery is tested against the downtime target and the gap between tested and required recovery is reported.
Can you state a number for coverage or effectiveness, and show the trend? - Governed
An owner runs scenario exercises, reviews objectives against plant change, and gaps drive funded investment.
Is there an accountable owner, a review cadence, and retained evidence?
Validation procedures
Until these pass, the practice is configured — not deployed.
- Test-restore a controller's logic/configuration and confirm the process resumes correctly.
- Verify a defined safe-state or manual fallback exists and is understood by operators.
- Confirm the last recovery exercise met — or exposed a gap against — the downtime target.
Evidence to retain
OT resiliency/contingency plan with downtime targets and fallbacks
Controller backup inventory and redundancy/spares list
Recovery-exercise reports and single-point-of-failure remediation
Test-restore results measured against the downtime target
Retaining these supports your own assurance and gives a reviewer something concrete to examine. It does not constitute an assessment or satisfy a contractual requirement on its own.
Operating metrics
| Metric | How it is calculated | Directional target |
|---|---|---|
| Controller backup coverage | Critical controllers with a current logic and configuration backup ÷ critical controllers. | 100%, refreshed after every change. |
| Tested recovery time | Elapsed time of the most recent restore test, compared with the agreed acceptable downtime. | Within the target; shortfalls documented and funded. |
| Unmitigated single points of failure | Critical components with neither redundancy nor an obtainable spare. | Zero, or an accepted risk with a named owner and a date. |
Targets are directional guidance for your own programme, not compliance thresholds.
Common failure modes
Backing up servers but not PLC logic and set points, redundancy that has never been failed-over to test it, and a recovery plan that assumes parts and expertise you cannot get quickly. Untested redundancy is a hope with a wiring diagram.
Two sized paths
Back up the logic and configuration for every controller you would struggle to rebuild, keep a copy somewhere that is not the plant floor, and restore one of them to prove it works. Then write down which spare parts have long lead times. That is most of the practical value.
Exercise recovery end-to-end for realistic failure and attack scenarios, measure against per-process downtime objectives, verify redundancy by actually failing over during planned windows, and let measured gaps drive the resiliency investment plan.
Tool categories
Categories, not recommendations. This site ranks no vendors and accepts no paid placement.
Backups age out the moment logic changes, so make a backup refresh part of the change record itself. Test failover and recovery in a lab or planned window with the process owner — an untested redundancy that fails during a real event is worse than a known gap.
Framework mappings
Independent mappings are aids to your own analysis — not authoritative equivalence, not coverage, and not a compliance determination. Read the caveat on every row before using it.
Why: The publication addresses recovery and restoration for control systems including configuration backups and degraded-mode operation.
What this does not claim: 800-82 is guidance, not a compliance standard.
Why: The guide is specifically about regular OT backups, their integration with change management, testing, and recovery exercises — the operational core of this practice.
What this does not claim: A quick-start guide, advisory by nature. It informs how backups and recovery testing are run; it imposes no requirement and does not cover redundancy, spares, or safe-state design.
Why: Recovery-plan execution and backup creation are the two outcomes this practice produces.
What this does not claim: The CSF describes outcomes, not testable controls.
Why: Control system backup and recovery/reconstitution are explicit system requirements in the availability family.
What this does not claim: IEC 62443 is a voluntary industrial standard; conformance is a separate, scoped exercise. The specific requirement identifiers are pending verification against the licensed normative text and should be treated as directional until that check completes.
Method: Read against the primary source text, then classified by relationship type and confidence. No automated mapping tool was used. How mappings are made →
NIST SP 800-171 relationships have their own requirement-level section below, with links into the full mapping experience.
Related NIST SP 800-171 requirements
Requirement-level relationships from the site’s independent 800-171 mapping. Each one names what it does and does not claim — a practice supports implementation of a requirement; it never satisfies one by itself. Expand a row for the rationale and caveat.
NIST SP 800-171 Rev. 2
3.8.9Backup CUI confidentialityPartial implementation supportLow
Why: The practice's off-device stores of controller logic, configuration, and set points are backup media, and where historian data, engineering workstation images, or project files in those stores contain CUI, protecting the copies advances this requirement.
What this does not claim: Directional only: most controller backups contain no CUI, so this mapping rarely applies. The practice's aim is availability — recovering production fast — and confidentiality protection of the backup store is not its default posture; where CUI is present, encryption and access limitation must be added deliberately, and adding them must not leave the plant unable to reach its own recovery files during an outage.
Review status: Technical review complete.
Open the 3.8.9 page →NIST SP 800-171 Rev. 3
03.08.09System Backup — Cryptographic ProtectionPartial implementation supportModerate
Why: System resiliency for OT leans on maintained configuration and system backups for critical production assets, and those backup stores fall inside this requirement's scope the moment they contain CUI — protecting them is part of the practice's recovery posture.
What this does not claim: The practice's center of gravity is availability — keeping production running and restoring it fast — while this requirement is about the confidentiality of what the backups contain. Encrypting OT configuration backups and restricting access to the repository are additional steps the resiliency work does not automatically include, and older OT backup tooling may not support encryption at all, forcing compensating protection of the storage location itself.
Review status: Pending NIST SME review.
Open the 03.08.09 page →Review status
| Technical review | Reviewed — pending SME sign-off |
|---|---|
| Editorial review | Reviewed |
| Reviewed by | inDirectIT practitioner review — CUI security and NIST SP 800-171 engineering |
| Last reviewed | |
| Official source verified | |
| Content version | 1.0 |
This guide has been reviewed by practitioners but is awaiting sign-off from a subject-matter expert in this specific domain. Treat the safety and change-control guidance as a floor, not a ceiling, and validate it against your own process and vendor requirements.