The short answer
Isolating a Compromised System
Isolation is the first action once a compromise is suspected, and it is done at network boundaries that were chosen and documented before the day: the office-to-control firewall, the DMZ links, the switch ports of individual machines, the remote site links, and the control-to-safety boundaries. The operator on shift may isolate the office connection and any single machine without waiting; the incident lead decides wider cuts. Machines are disconnected from the network, not powered off, so that memory and logs are preserved for investigation, unless they are actively destroying something. The process continues on the controllers with local control and, where needed, on the manual operation procedures, which is why the plant was designed to run without the SCADA. If the compromise is still spreading, isolation widens along the zone structure until it stops, and every action is logged with the time.
Key points
- Decide the isolation points before the incident and draw them on the network diagram; the moment of the incident is not the time to design.
- The operator may cut the office connection and a single machine at once; wider cuts are the incident lead call, made quickly.
- Disconnect the network, not the power, unless the machine is actively causing harm; evidence lives in memory.
- The controllers keep running the process; local control and manual procedures are the reason isolation is possible.
- Widen along the zones: machine, segment, zone, site, everything from the office.
- Log every cut with the time and who; the timeline is evidence and the basis for reconnection.
Before the day
Isolation works when the cuts have been planned. The network drawing marks the isolation points: the firewall interface to the office, the DMZ links, the uplink of each switch, the port of each server and workstation, the radio and cellular links to remote sites, and the boundaries between the control zone and any safety or protection system. For each, the drawing says how to cut it, physically pulling a cable or disabling a port or an interface, where the credentials and the console cable are, and what stops working when it is cut. The list is printed, kept in the control room and the safe, and rehearsed in the annual exercise. A plant that has practiced running for a day with the office link and the DMZ cut has learned what breaks, and fixed it, before an incident makes it urgent.
Who decides
| Action | Who may take it | When |
|---|---|---|
| Cut the office-to-control firewall link | The operator on shift | On any credible sign of compromise, without waiting for anyone |
| Disconnect a single machine from the network | The operator on shift, or any engineer | When that machine is behaving abnormally: unexpected screens, encryption notices, unknown processes, control actions nobody took |
| Cut the DMZ links and remote access | The incident lead, or the operator if the lead cannot be reached in minutes | When the compromise is confirmed or the source is unknown |
| Isolate a zone or a site | The incident lead | When the compromise is spreading or a zone is the suspected source |
| Isolate the control zone from every other network | The incident lead | When the extent is unknown; this is the default when in doubt |
| Stop a controller or a process | The operator, by the manual operation procedures | Only when the controller itself is compromised and the process is unsafe under its control |
How to cut
- At the firewall: disable the interface or the rules for the office and the DMZ; a firewall that can be managed only through the office link has a console cable and a local login for this purpose.
- At the switch: disable the port of the affected machine, or the uplink of the affected segment, from a console or a management session on the control side.
- Physically: pull the cable at the machine or the patch panel; slower to undo but certain, and it needs no credentials.
- Remote access: disable the VPN and the jump host accounts; vendor connections are cut with the DMZ.
- Remote sites: leave the site links up unless a site is the source or the target; the sites run on local control either way.
- Wireless: disable any wireless access on the control network.
Keeping the process running
The isolation of the SCADA does not stop the controllers; they run their logic on their own I/O. Operators lose the displays, the alarms, and the remote control, and they fall back to the manual operation procedures: the local HMI panels, the physical gauges and levels, rounds of the sites, and hand control at the panels where required. The procedures say who goes where, how often, what to record on paper, and what conditions require a process to be shut down. If a controller itself is suspected, the operator takes the process to hand control and the controller is isolated and left running for the investigators. The incident plan names the people, the phone tree, and the shift arrangement for a manual operation that may last days.
Widening
- 1
One machine
Disconnect it; watch whether the abnormal behavior appears elsewhere. Read the logs on the collector for what it talked to.
- 2
The segment
If other machines on the same VLAN or switch show signs, or the logs show lateral movement, cut the segment uplink.
- 3
The zone
If the supervisory zone is affected, cut it from the control zone; the controllers continue without it.
- 4
Everything from outside
Office, DMZ, remote access, vendor links, wireless. This is the default when the extent is unknown, and it is not a failure to reach for it first.
- 5
Sites
Cut any remote site that is the source or that shows signs; the rest stay connected and are watched.
- 6
Hold
Nothing is reconnected until the investigation says the threat is understood and the recovery plan says the reconnected part is clean.
Preserve and record
- Every action with the time and the person, on paper in the control room and in the incident log.
- The logs on the collector, the firewall, and the switches, copied to media before anything is rebuilt.
- Images of the affected machines taken by someone who knows how, before they are wiped; if nobody in the utility does, the machines are left disconnected and running until the responders arrive.
- Screens photographed, ransom notes saved, and unusual controller behavior described.
- The state of the process at the time of isolation: levels, pressures, which pumps were running, what the operators did.
Frequently asked questions
- What if isolating the office link stops something we need?
- Then the design has a dependency that has to be removed before the incident, which the exercise finds. The control system runs without the office: time from a local source, authentication from a control-side directory or local accounts, no historian or database on the office side. Anything that breaks when the office is cut is on the list to fix.
- Should we shut down the plant to be safe?
- Rarely. The controllers are running the process; shutting down creates its own hazards and a public health event. Isolate the network, keep the process running on local and manual control, and stop only a process whose controller is compromised and cannot be operated safely by hand. The manual operation procedures are what make that judgment possible.
- How long can we stay isolated?
- As long as the manual operation can be sustained with the staff available, which the procedures state: days for most utilities, with rounds and paper logs. The recovery plan works toward reconnection in stages, and the isolated state is not lifted to relieve inconvenience.
- Who do we call?
- The incident plan lists them: the utility incident lead, the integrator, the incident response firm if one is retained, the federal cybersecurity agency and law enforcement, the state regulator, and the sector information sharing center. The calls are made while the isolation is being done, not after, and the numbers are on paper.
Related topics
- OT Incident Response PlanThe plan for the day the control system is not trusted: who decides, the triggers, the first hour, isolating without stopping treatment, running by hand, preserving evidence, who to notify and when, recovery from backups, and the exercise that makes it real.
- Manual Operation ProceduresWritten procedures for running the plant and the remote sites without the SCADA or the controllers: what to read and set by hand, how often to check, what to log, chemical dosing from a flow reading, when to stop, staffing, and the drill that proves them.
- Recovery and RestorationBringing a control system back after a compromise without reintroducing it: controllers verified against baselines first, servers rebuilt from clean media and project exports, every credential changed, hardening, staged reconnection, and end-to-end tests.
- Zones and ConduitsThe IEC 62443 way to segment a control system: grouping assets into zones with a shared security level, inventorying every conduit between them, and turning the drawing into firewall rules. With a worked water utility example.
- Reporting RequirementsWho a water or wastewater utility must tell about a cyber incident, how fast, and what to say: the federal critical infrastructure reporting rule, state laws, the drinking water regulator, public notification, law enforcement, the sector center, and insurers.
- Industrial DMZ DesignThe buffer zone between the business network and the control system: what goes in it, the no-direct-path rule, push-not-pull data flows, the firewall pair, and the services a utility actually needs to place there.
Direct contact
Have a controls question?
Reach Eric Sullivan directly about anything on this site, a controls or automation topic, or one of his personal projects.