The short answer
Alarms Nobody Answers
An alarm nobody answers is an alarm that should not exist in its current form. Every alarm must require a response from a person; if it does not, it is an event to log, not an alarm. A plant gets to hundreds of alarms a day through defaults, missing deadbands and delays, duplicated conditions, and alarms that were never removed after the reason for them went away. The way back is a rationalization: go through the list, keep the alarms that have a response, fix their limits and delays, and turn the rest into logged events.
Key points
- An alarm is a call for a person to act; anything else is an event and belongs in the log, not on the banner.
- ISA-18.2 gives the benchmarks: around 150 alarms per operator per day is very likely acceptable, 300 is the most that is manageable, and more than ten in ten minutes is a flood.
- Most nuisance alarms come from limits with no deadband, no delay, or no thought, and from conditions alarmed twice.
- A standing alarm that has been active for weeks is either a broken thing nobody fixed or an alarm that should not exist.
- Rationalization is a meeting with the operators, the list, and a rule for every entry; a small plant can do it in a day.
The on-call phone at a small wastewater plant rang so often that the operators had an arrangement: whoever had the phone silenced it at night and looked at the screen in the morning. It was a sensible response to a system that called about a communication blip every time the radio faded, a low chlorine flow every time a pump changed over, and a high level in a tank that had read high since the transmitter was replaced two years earlier. The night the wet well actually overflowed, the phone rang, and the phone was silenced, and nobody could honestly say the operator had done anything different from every other night.
The alarm system had not failed that night. It had failed slowly over years, one reasonable alarm at a time, until it carried so much noise that the signal could not get through. The word for the fix is rationalization, and it sounds like a project. At a small plant it is a day.
What an alarm is for
ISA-18.2, the standard for alarm management, defines an alarm as an indication that requires a response from an operator. That definition is the whole discipline. A pump starting is not an alarm; it is an event. A communication link dropping for ten seconds and returning is an event. A tank that reads high because the transmitter is wrong is a maintenance item, and the alarm for it should be shelved until the transmitter is fixed. When every alarm on the list has a response someone can name, the list is short and the operator answers it. When the list is padded with things that need no response, the operator learns the only lesson available, which is that the alarms do not matter.
How a plant gets to eight hundred a day
- Defaults
- Every tag came with high and low alarms because the SCADA template had them. Nobody decided that a filter effluent turbidity of 0.31 NTU needed a phone call; the template did.
- No deadband
- A level that crosses 80 percent and hovers there generates an alarm every time it wiggles. A deadband of a few percent turns fifty alarms into one.
- No delay
- A flow that drops during a pump changeover alarms for the ten seconds it takes the next pump to come up. A thirty second on-delay removes every one of those.
- Duplicates
- The same condition alarms in the controller, in SCADA, and again as a derived alarm, so one event is three entries and three phone calls.
- Consequences
- One root event, such as a power dip, trips forty alarms as everything restarts. Without suppression of the consequences, the operator has to read forty to find the one.
- Leftovers
- The alarm added during a startup problem in 2016 is still there, because nobody owns removing alarms.
The numbers to hold the system against
ISA-18.2 and the EEMUA 191 guidance behind it give benchmarks that a small plant can use directly. Around 150 alarms per operator per day, about six an hour, is very likely to be acceptable. Around 300 a day is the maximum that is considered manageable. More than ten alarms in any ten minute period is a flood, and a plant that floods regularly has an alarm system that stops working when it is most needed. The other number is the standing alarm count: how many alarms are active at a calm moment on an ordinary day. On a healthy system it is close to zero. On the plant with the silenced phone it was thirty one.
| Measure | Healthy | Needs work | What it usually means |
|---|---|---|---|
| Alarms per day, one operator | Under 150 | Over 300 | Limits, deadbands, and delays never set |
| Peak in any ten minutes | Under 10 | Over 10 regularly | Consequence alarms not suppressed |
| Standing alarms on a calm day | Near zero | Ten or more | Broken instruments and alarms with no response |
| Share from the top ten alarms | Small | Half or more of all alarms | A few nuisance alarms are the whole problem |
| Alarms with a priority other than the default | Most | Few | Priority was never assigned, so it means nothing |
The top ten
Before any meeting, pull the alarm history for a month and count by alarm. In nearly every plant the top ten alarms are half or more of the total, and they are the same handful of nuisance conditions repeating. Fixing those ten, usually with a deadband, a delay, or a decision that the condition is an event, cuts the load in half before anyone rationalizes anything. It is the highest return hour in the whole exercise, and it makes the meeting shorter because the operators see the system get quieter immediately.
The rationalization day
- 1
Bring the right people
The operators who carry the phone, whoever maintains SCADA, and someone who can change the controller. Not a committee; the people who answer the alarms and the people who can change them.
- 2
Take the list in order of frequency
Start with the alarms that fire most. For each one ask the only question that matters: what does the operator do when this goes off?
- 3
Give every kept alarm its settings
A limit that someone chose, a deadband, an on-delay, and a priority that reflects how fast the response has to happen and what happens if it does not.
- 4
Turn the rest into events
Anything with no response is logged, trended if useful, and removed from the banner and the phone.
- 5
Handle consequences
For a power dip, a communication loss, or a pump trip, decide which one alarm the operator needs and suppress the rest while the root condition is active.
- 6
Write the philosophy
Two pages: the definition, the priorities and their meaning, the response time expected for each, and the rule that nobody adds an alarm without those settings.
Keeping it quiet
A rationalized system drifts back unless something stops it. The stop is a rule that every new alarm arrives with a response, a limit, a deadband, a delay, and a priority, written down, and a monthly look at the top ten and the standing list. The look takes ten minutes. When a new nuisance appears, it is fixed the month it appears rather than the year the phone gets silenced. Shelving, a timed suppression of an alarm the operator knows about, gives the operator a way to quiet a broken instrument without deleting the alarm, and the shelved list is the maintenance list.
What changed at the plant
The plant with the silenced phone went from about seven hundred alarms a day to under sixty in a week. The top ten had been eleven radio communication alarms, a chlorine flow alarm with no delay, and the tank that read high. The radio alarm became one alarm per site with a two minute delay, the chlorine flow got thirty seconds, the transmitter got replaced. The phone rings a few times a week now. It gets answered.
Frequently asked questions
- Is it safe to remove alarms? What if we need one later?
- An alarm that is not answered is not protecting anything. Removing it from the banner and keeping the condition as a logged event loses nothing, and the event history is there if the condition turns out to matter. Add it back with proper settings if it does.
- Where should alarm limits be set, the controller or SCADA?
- The condition is best evaluated in the controller, so it survives a SCADA outage and drives local logic. SCADA presents, prioritizes, records, and notifies. Whichever side holds the limit, only one should, and the other reads it.
- We have one operator for six sites. Do the benchmarks still apply?
- They apply per operator, which means the six sites together should be inside them. That usually means each site has to be quieter than a single plant would be, and it makes consequence suppression and delays more important, not less.
- Do we need alarm management software to do this?
- No. The alarm history in SCADA and a spreadsheet are enough for the counts, and the rationalization is a conversation. Software helps at large plants with thousands of alarms; a small plant needs the discipline more than the tool.
Related topics
- The Alarm PhilosophyThe document that decides what is allowed to be an alarm: the definition, the criteria, the priorities and their meaning, the performance targets, the handling rules, and who owns it. Under ISA-18.2 it comes first.
- Alarm RationalizationThe meeting where every alarm earns its place: who attends, what is decided for each alarm, how to document it in a master alarm database, and how to run it at a utility that cannot spare a week.
- Alarm FloodsWhat an alarm flood is, why it happens at exactly the wrong moment, the ISA-18.2 rate targets, and the design measures that keep a power failure or a communication loss from burying the one alarm that matters.
- ISA-18.2 Alarm ManagementThe lifecycle standard for alarm systems, the rate targets that define a workable system, and why alarm rationalization is the step nobody wants to fund.
- Alarm ShelvingThe operator tool for temporarily silencing a nuisance alarm: what shelving is and is not, the rules a philosophy sets, who may shelve what and for how long, automatic unshelving, the shelved list, the record, and how to implement it on any platform.
- How to Configure AlarmsConfigure alarms in the SCADA system from the alarm list: set the source, controller-generated or evaluated in SCADA, the limits and deadband and delay, and the notification; then test each alarm end to end and review the alarm rate after the first weeks.
Direct contact
Have a controls question?
Reach Eric Sullivan directly about anything on this site, a controls or automation topic, or one of his personal projects.