The short answer
Alarm Rationalization
Alarm rationalization is the systematic review of every configured alarm against the alarm philosophy. For each one, a small team decides whether it is a real alarm, what its cause and consequence are, what the operator is expected to do, what its setpoint and priority should be, and records the result in a master alarm database. It is how a system gets from thousands of nuisance alarms to a set operators trust.
Key points
- Every alarm must have a defined operator action. If there is nothing to do, it is not an alarm.
- The team is an operator, a process or controls engineer, and a facilitator. Two hours at a time, not two weeks.
- Record cause, consequence, action, time to respond, setpoint, priority, and classification for every alarm.
- The master alarm database is the record. SCADA configuration must match it, not the other way around.
- Start with the ten most frequent alarms. That is where the noise is.
What rationalization is
Rationalization is the stage in the ISA-18.2 lifecycle where the alarm philosophy is applied to actual alarms, one at a time. It answers, for every configured alarm, whether the alarm should exist and, if so, what it means and what it demands. Without it, an alarm system is whatever accumulated: the defaults every tag came with, the alarms added after each incident, the ones a contractor thought were sensible in 2009.
The output is a master alarm database, a record of every alarm and the decisions made about it. That record is what makes the alarm system maintainable. When someone asks why a setpoint is where it is, or whether an alarm can be removed, the answer is written down.
The questions asked of each alarm
| Question | What is recorded | Why it matters |
|---|---|---|
| Is it an alarm? | Whether it meets the philosophy definition: an abnormal condition requiring operator action | Anything with no operator action is an event or a maintenance notification and leaves the alarm system. |
| What causes it? | The likely process or equipment causes | Guides the operator response and shows whether the alarm duplicates another. |
| What is the consequence? | What happens if no one acts, in the philosophy categories | Drives the priority and justifies the alarm. |
| What must the operator do? | The specific corrective action | The single most important entry. An alarm the operator can do nothing about is not an alarm. |
| How long do they have? | Time available to respond before the consequence | Drives the priority together with the consequence. |
| What is the setpoint? | The value and the basis for it: a limit, a permit, a curve, an operating envelope | A setpoint with a basis survives; one without is changed whenever it annoys. |
| What priority? | From the matrix in the philosophy | Consistency across the system. |
| What class? | Safety, environmental, regulatory, or general; determines testing and change control | A regulatory alarm cannot be changed by an operator at a keyboard. |
| Deadband, on-delay, off-delay? | Values that prevent chattering | Most nuisance alarms are cured here. |
| Suppression or shelving rules? | When the alarm is expected and should not annunciate | A pump-stopped alarm on a pump that was commanded to stop is designed suppression. |
Who is in the room
Rationalization is a small-group activity. The essential people are an experienced operator who knows what the alarm looks like in practice and what they actually do about it, an engineer who knows the process and the control system, and a facilitator who keeps the pace and writes things down. A maintenance representative is valuable for equipment alarms. Managers are welcome to set the philosophy and to approve the result; they slow the sessions down if they attend them.
Sessions run about two hours. Longer than that, decisions get worse. A practiced team gets through 20 to 40 alarms an hour once the philosophy is settled, more where alarms are similar, such as the same set on each of thirty lift stations, which can be rationalized as a template and then checked for exceptions.
Running it at a small utility
A utility with two operators and no engineer on staff cannot follow the full program as written for a refinery. It can still rationalize. The shortcut that works is to start with the alarm history.
- 1
Pull the last three months of alarms
Count occurrences per alarm tag. Ten tags usually account for half the total. Those ten are the first session.
- 2
Fix the bad actors
Most are chattering on a noisy signal, standing because a piece of equipment is out of service, or duplicating another alarm. Deadband, delays, and suppression rules cure most; a few are deleted.
- 3
Rationalize the high-priority alarms
Every alarm currently marked high gets the full set of questions. This is where the priority distribution is fixed.
- 4
Template the repeated sites
Rationalize one lift station completely, then apply the result to the others and review only the differences.
- 5
Work through the rest by area
One process area per session, on a schedule. It takes months, and the alarm system improves every session.
- 6
Keep the database current
Every new alarm goes through the questions before it is configured. Every change to a setpoint or a priority updates the record. The database is the configuration authority.
The master alarm database
The database can be a spreadsheet, a table in the SCADA system, or a commercial alarm management package. What matters is that it holds one row per alarm with the answers above, that it is under change control, and that it is periodically compared with what is actually configured in SCADA. Drift between the two is normal and is found by an audit, not by an incident. Commercial tools automate the comparison and push approved settings to the SCADA system; a spreadsheet and a quarterly check achieve the same thing at small scale.
After rationalization
Rationalization is not the end of the lifecycle. The rationalized alarms are configured, operators are trained on the changes, and the system is monitored: alarm rate per operator per hour, the ten most frequent alarms, standing alarms, the priority distribution. Those metrics identify the alarms that need to go back through the questions. ISA-18.2 gives the targets, and the alarm floods page covers what to look for when the rate spikes.
Frequently asked questions
- How long does rationalization take?
- Rough planning numbers are 20 to 40 alarms per hour of session time once the philosophy is settled, with templated sites going much faster. A plant with 2,000 configured alarms is a few hundred hours spread over several months. A utility with a SCADA system built from defaults often finds that a third of the alarms are removed outright.
- Do we need an alarm philosophy first?
- Yes, at least a short one: the definition of an alarm, the consequence categories, the priority matrix, and the rules for classification. Rationalizing without it produces inconsistent decisions, and the first sessions become arguments about principles instead of alarms. A philosophy for a utility can be ten pages.
- What do we do with alarms the operator cannot act on but management wants?
- Route them somewhere other than the operator alarm list: a maintenance notification queue, an email report, an event log. The information is kept; the operator is not interrupted by it. If management insists it be an operator alarm, the operator action must be defined, and the action cannot be to call management.
- Should we buy software for this?
- Alarm management software helps with the analysis, the database, and enforcement, and pays for itself at a plant with thousands of alarms. At a small utility, a spreadsheet, the SCADA alarm history export, and a disciplined process do the same job. Buy the software when the spreadsheet stops being maintainable.
Related topics
- ISA-18.2 Alarm ManagementThe lifecycle standard for alarm systems, the rate targets that define a workable system, and why alarm rationalization is the step nobody wants to fund.
- The Alarm PhilosophyThe document that decides what is allowed to be an alarm: the definition, the criteria, the priorities and their meaning, the performance targets, the handling rules, and who owns it. Under ISA-18.2 it comes first.
- Alarm PriorityHow to assign alarm priorities that operators trust: a consequence-and-time matrix, three or four levels, the target distribution from ISA-18.2 and EEMUA 191, and the mistakes that make every alarm high.
- Alarm FloodsWhat an alarm flood is, why it happens at exactly the wrong moment, the ISA-18.2 rate targets, and the design measures that keep a power failure or a communication loss from burying the one alarm that matters.
- Alarm Notification and CalloutGetting the alarm to the person on call when no one is watching the screen: notification paths, escalation, acknowledgment from the field, which alarms qualify, and the failure modes that leave a station in high level with nobody paged.
- Alarm ShelvingThe operator tool for temporarily silencing a nuisance alarm: what shelving is and is not, the rules a philosophy sets, who may shelve what and for how long, automatic unshelving, the shelved list, the record, and how to implement it on any platform.
Direct contact
Have a controls question?
Reach Eric Sullivan directly about anything on this site, a controls or automation topic, or one of his personal projects.