The short answer
Recovery and Restoration
Recovery starts once the intrusion is contained and understood, and it proceeds from the process outward: the controllers are verified against their configuration baselines or reloaded from the offline copy and watched in hand before they are trusted in automatic; the servers are rebuilt from clean installation media and the project exports, not from images taken while the intruder was present, with the historian archives restored from backups that predate the compromise where possible; every credential in the system is changed and the vault is rebuilt; the hardening baseline and all pending patches are applied before any machine is reconnected; the network is reconnected zone by zone with monitoring for any sign that the intruder remains; alarms and loops are tested end to end; processes are returned from manual to automatic one at a time with an operator watching; and the incident is closed with a review that changes the design, the procedures, and the plan.
Key points
- Order: controllers verified, then servers rebuilt, then network reconnected, then processes returned to automatic.
- Rebuild from clean media and project exports; an image taken during the intrusion restores the intruder.
- Every credential changes: accounts, service passwords, device passwords, keys, certificates.
- Harden and patch before reconnecting; the vulnerability that was used is still there otherwise.
- Reconnect in stages with monitoring, and be ready to isolate again.
- Test loops and alarms before trusting them; return processes one at a time with eyes on them.
When to start
Recovery begins when the investigation has established how the intruder got in and what they touched, or has at least established the boundaries of what cannot be trusted. Starting earlier rebuilds systems that are reinfected from a machine nobody looked at. While the investigation runs, the plant operates by the manual procedures, and the recovery team prepares: clean installation media verified, the offline backups retrieved and checked, spare hardware or virtual machines staged, credentials generated, and the sequence written down. The pressure to reconnect is real and it is resisted until the plan says a part is clean.
The sequence
- 1
Controllers first
For each controller and remote unit: compare the running program against the baseline with the vendor tools. A match, and a review of the data and setpoints, means it can be trusted. A difference means the program is reloaded from the offline copy, the firmware is verified or reloaded, the controller credentials are changed, and the process is watched in hand through a cycle before the controller is returned to automatic.
- 2
Servers from clean media
Each server rebuilt on a fresh virtual machine or wiped hardware from verified installation media, the operating system hardened to the baseline and patched, the application installed at a supported version, and the project imported from an export that predates the compromise or has been reviewed. Images taken during the intrusion are kept for evidence and not restored.
- 3
Data
The historian archives restored from backups, with the period of the intrusion treated as suspect and marked; the alarm journal and audit database restored to the last clean backup and the incident period reconstructed from the collector logs.
- 4
Credentials
Every account password, every service account, every device password, every key and certificate, and the vault itself rebuilt with new master credentials. Any account the intruder could have created is found and removed.
- 5
Network
Firewall and switch configurations restored from baseline and reviewed rule by rule; the rule or the path that was used is closed; monitoring in place at the boundaries before anything is reconnected.
- 6
Reconnect in stages
Control zone first, with the rebuilt servers polling the verified controllers and the operators watching; then the supervisory zone; then the DMZ with the historian replica; then remote access, last, with multi-factor authentication confirmed. At each stage, a pause and a look at the monitoring for anything unexpected.
- 7
Test
Loop checks on critical signals, alarm tests end to end including notification, failover of the redundant pair, and a walk of the displays against the plant.
- 8
Return to automatic
Process by process, with the manual log closed for each, the setpoints checked, and an operator watching the first cycles.
- 9
Close
The incident declared over by the incident lead with management, the reports filed, the evidence retained, and the review scheduled.
Rebuild, not restore
The instinct after an intrusion is to restore the server image from last night and get the screens back. The image from last night contains whatever the intruder installed last week, and restoring it returns the system to them. The clean path is longer: install from media whose hashes have been verified, apply the hardening baseline, patch, and import the project from an export that has been reviewed, or from before the earliest sign of intrusion. The project export is data and configuration rather than executable code, which is why it can be trusted where the image cannot, and it is why keeping project exports as well as images matters. The rebuild also produces a system that is hardened and patched, which the original may not have been.
Testing before trust
- Every controller: program compare, firmware version, mode, and a watched cycle of the process.
- Critical loops: a loop check from the instrument to the display, and a command from the display to the equipment.
- Alarms: the highest priority alarms forced and followed to the display, the journal, and the phone.
- Communications: every remote site polling, with quality good and timestamps correct.
- Redundancy: a failover of the server pair and a check of alarm states afterward.
- Accounts: every login tested for its role; no account that was not created by the rebuild.
- Backups: the first backup of the rebuilt system taken, and an offline copy made, before the recovery is declared complete.
After
The review is held within weeks, with everyone who was involved, and it produces changes: to the network design that allowed the entry or the spread, to the procedures that were unclear, to the backups that were missing something, to the monitoring that did not see it, to the plan that did not name a person. The incident report goes to management and, as required, to the regulators and the sector organizations, and the lessons are shared where the utility can share them. The recovery time, from isolation to the last process in automatic, is measured and becomes the new planning number.
Frequently asked questions
- How long does recovery take?
- Days to weeks for a utility with tested backups, clean media, and a plan; months for one without. The controllers and the critical processes come back first, within days if the offline copies exist; the full SCADA with history and remote access takes longer. Manual operation covers the gap, which is why its sustainable duration matters.
- Can we trust the controllers if only the servers were hit?
- Not without checking. An intruder on the SCADA server had a path to every controller it polled. Compare every controller program against its baseline; a match is evidence, and a difference is a finding. Controllers that match and whose data looks right can be trusted while the servers are rebuilt.
- Should we pay a ransom to speed recovery?
- That is a decision for management with counsel, law enforcement, and the insurer, and the federal guidance advises against it. A utility with offline backups and a tested restore has a recovery path that does not depend on the intruder, and the decision is easier. Recovery from backups is the plan; the ransom is not.
- What if the offline backups are old?
- Then the changes since the backup are re-created from the change records and the documentation, which is another reason to keep both. An old clean backup plus a list of changes is recoverable; a current backup that contains the intruder is not.
Related topics
- OT Incident Response PlanThe plan for the day the control system is not trusted: who decides, the triggers, the first hour, isolating without stopping treatment, running by hand, preserving evidence, who to notify and when, recovery from backups, and the exercise that makes it real.
- Isolating a Compromised SystemCutting a compromised machine, segment, or site off from the control system without stopping the process: who may decide, isolation points planned in advance, disconnecting the network rather than the power, running on manual control, and preserving evidence.
- Offline CopiesThe backup copy that survives ransomware: why anything reachable over the network will be encrypted too, what offline and immutable mean, options from rotated drives to immutable storage, what goes in the offline set, and the restore test that uses it.
- Backup TestingWhy backups that have never been restored fail when needed, and how to find out: spot restores, full restores of a server to an isolated machine, controller program restores to a spare, archive reads, timing each one, and the findings and record that follow.
- Configuration BaselinesA known-good record of how every part of the control system is configured, captured at acceptance and kept current through change management: what to baseline, how to capture and store it, how to compare on a schedule, and what an unexplained difference means.
- Reporting RequirementsWho a water or wastewater utility must tell about a cyber incident, how fast, and what to say: the federal critical infrastructure reporting rule, state laws, the drinking water regulator, public notification, law enforcement, the sector center, and insurers.
Direct contact
Have a controls question?
Reach Eric Sullivan directly about anything on this site, a controls or automation topic, or one of his personal projects.