Skip to content

Don’t be alarmed by Nuisance Alarms !

Bart was excited to be sitting in front of the operator screens at the power plant. It was his first day as a junior operator. As he checked the monitors, two alarms pop-up on the screen. He turned to Homer, one of the senior operators and said, “Hey Homer, two alarms have just appeared on the alarm list for Unit 700. Could you have a look?” Homer navigated to the unit on his monitors, looked at the alarms and said, “Acknowledge them”. “Don’t we need to take some action?” asked Bart. “Nope” replied Homer, “Those are just a nuisance. They pop up a few times every hour.”

The ideal scenario for an operating plant is to have every control system alarm indicate a malfunction or abnormal condition that requires an operator’s action. These alarms are relevant, unique, timely, and prioritised with clear diagnostic information and advisory action. 

In reality, several alarms are irrelevant or annunciate repeatedly even after being cleared, often not indicating an abnormal process condition and not requiring any operator action. These are called nuisance alarms. 

Nuisance alarms pose a risk to the successful operation of the plant because they overload operators with nonessential information and lead to a loss of confidence in using the alarm system to detect abnormal events. As little as a 25% false-alarm rate is enough for operators to stop relying on the alarm system.

When a LOPA is carried out, credit is taken for having an alarm system in place. Therefore, when the alarm system is not properly managed this can have severe consequences. As an example,  one of the causes of the explosion at the Milford Haven Refinery was the poor design of the alarm systems and the high number of alarms that distracted the control room operators from critical safety alarms. Details of this report can be found here [1]. 

Why do we have nuisance alarms?

The ISA 18.2 standard defines an alarm as an audible and/or visible means of indicating to the operator an equipment malfunction, process deviation, or abnormal condition requiring a response.

One of the most important principles of Alarm Management is that an alarm requires a response. Hence, if the operator does not need to take action within a certain time, then there should not be an alarm. Alarms generated at the right time with a defined action for the control room operator can help minimise plant upsets, prevent asset damage and avoid safety and environmental incidents. Alarms must exist for the benefit of the operators and do not replace the surveillance of qualified operators. 

The ISA 18.2 Alarm Management standard provides a framework for the successful design, implementation, operation, and management of alarm systems in a process plant. This standard is organised around the alarm management lifecycle, which contains ten stages with key alarm management activities executed in each stage of the lifecycle. The outputs of each stage of the lifecycle are the inputs for the activities of the next stage.

alarm-lifecycle
ISA 18.2 Alarm Management Lifecycle

Depending on the status of the operating plant, the lifecycle can begin at three potential starting points: Philosophy, Monitoring & Assessment, and Audit. Philosophy is the typical starting point for a new facility while monitoring & assessment or performing an audit to benchmark the alarm system may be the starting point for an existing facility. 

In the past, control rooms had hardwired panel board alarms for an operator to monitor. Installing a new alarm required planning, additional material costs and resources as the alarms were hardwired to lamps on display panels. The inclusion of a new alarm to the system involved panel design, additional IO and the associated installation and panel wiring activities to connect the alarm to the physical display panel board. 

Modern-day control rooms use computers for HMIs and process controllers with much larger memory capacities. Adding new alarms has become very easy and cheap, so engineers added alarms everywhere and on everything, without any thought or consideration to situational awareness and human factors. It was better to be safe than sorry. 

Without alarm philosophy, the control system vendors often ended up leaving the default alarm setting enabled for every tag, resulting in a possible 10 or more different alarms for a tag like HH, H, L, LL, Deviation, Discrepancy, Rate of change, Comms fail, Out of Range, Open circuit, Burn out, Over-range and Under-range alarms. Package vendors can also be provided with their own set of alarms and some of them may not be fully aligned with the plant alarm philosophy, if there is any.

As a result, the number of configured alarms per operator station has grown exponentially from an average of 100 in the 1960s to greater than 3000 in the early 2000s. This impacts the situational awareness of operators, which is the operators’ ability to understand the operating conditions of the plant. With too many alarms, the operator is overloaded with information, affecting the operator’s ability to detect, diagnose and respond to genuine alarms generated by the system.

The ISA 18.2 standard includes recommended system performance metrics for an alarm system. If the measured performance of an alarm system exceeds the prescribed values, the reliability of the operator’s response to an alarm will be compromised and this impacts the risk reduction provided by an alarm system as an independent protection layer.

alarm-metrics
ISA 18.2 Alarm Performance Metrics

Enter Alarm Rationalisation, a key phase of the Alarm Management lifecycle, where a situation or event is evaluated to determine if it qualifies as an alarm and to ensure that it meets the criteria for being an alarm as defined in the alarm philosophy. Rationalisation should result in a minimum set of alarms required to keep the process safe and within normal operating conditions. Each alarm is also given a priority, which is typically determined based on the severity of the potential consequences of not taking corrective action and the time to respond to the alarm.

Reviewing each alarm and justifying it, ensuring it is consistent with the overall philosophy, operating conditions and risk assessments is time-consuming. Getting executive buy-in and finding resources with the right expertise are common challenges that projects face when implementing alarm management.

Therefore, for an operating plant, alarm rationalisation can take several months. Before alarm rationalisation, HAZOPs and LOPAs need to be performed and the plant design should be completed so that alarms can be structured and distributed based on plant functional units for example process units, utility and service units etc. In addition, there are different operating conditions – start-up, partial operation, normal operation, normal shut-down, ESD and maintenance. 

In such scenarios, a quick win can be the identification and resolution of nuisance alarms in the operating plant. Remember that nuisance alarms are alarms that repeat frequently and unnecessarily or alarms that do not return to a normal state after the operator has taken the correct response. While most nuisance alarms are false alarms and may themselves be harmless, depending on how much the operator is overloaded by these alarms, it may cause the operators to ignore real high-priority alarms that need their attention, leading to a potentially dangerous scenario. This is explained in an interesting YouTube video created by Exida.

Nuisance alarms can be identified by examining the alarms generated on the alarm system and noting down the bad actors. When the most frequently occurring alarms are observed in an alarm system, 10 – 20 tags are usually responsible for creating nearly 50% – 80% of the alarms. These are the bad actors. It is not uncommon to find that most bad actor alarms are nuisance alarms because they are generated so frequently. Measuring the alarm system performance against the ISA 18.2 Alarm Performance Metrics can help verify the extent of the problem with nuisance alarms.

nuisance-alarms
Types of Nuisance Alarms

Types of Nuisance Alarms

  • Chattering Alarms – These are alarms that repeatedly transition between the alarm state and normal state in a short time frame, typically three or more times per minute. The alarm clears itself without any action from the operator but leads to the creation of many events cluttering the alarm summary and process graphic displays on the HMI, distracting the operators. 
  • Fleeting Alarms – These are short-duration alarms like chattering alarms, the difference being that they do not immediately repeat. 
  • Stale Alarms – These alarms remain in alarm state for an extended period, usually greater than 24 hours, with no operator action and do not clear after an action has been taken by the operator. The presence of stale alarms can lead to reduced operator effectiveness, clogging the HMI displays and making it more difficult to detect new alarms. Some alarms remain in the alarm state continuously for days, weeks, or months and provide little valuable information to the operators. A typical stale alarm example is an alarm that is generated because the equipment is in a different mode of operation. For example, a low discharge pressure alarm for a pump when the pump has been stopped, as part of normal operation. 
  • Redundant Alarms – These are alarms that duplicate other alarms that have the same operator response. Such alarms are usually associated with redundant equipment or redundant instrumentation. They can also include alarms where a second or third alarm is generated as a cascade from the condition that triggered the first alarm, with the other alarms requiring the same operator action as the first alarm. Another example of redundant alarms is having a High High alarm in addition to a High alarm, with the same operator response for both alarms. The High High alarm should not be set up just to account for the operator missing the High alarm unless it requires the operator to take a different action. 

As discussed earlier, an alarm philosophy and alarm rationalisation can help eliminate alarms that do not require any operator action. Fixing these bad actors can reduce the number of alarms generated by a significant percentage. While redundant alarms usually require alarm rationalization to eliminate the alarms, some of these nuisance alarms might be alarms that are recommended as part of the alarm rationalization. These nuisance alarms can often be fixed by properly configuring or conditioning the alarm. 

How can we manage nuisance alarms?

Nuisance alarms can be managed by Alarm Conditioning or Alarm Suppression techniques. 

nuisance-alarm-mngt
Nuisance Alarm Management Techniques

ALARM LIMIT SETPOINT – If the alarm limit is not set correctly, this can lead to stale alarms. The alarm limit set point should be set relative to the normal operating envelope of the process and relative to the consequence threshold. The consequence results when no operator action is taken or an incorrect or insufficient action is taken or the action is not completed within the allowable response time. The process condition limits at which the consequence begins to occur is the consequence threshold.  

DEADBAND – This is the change in the value of the signal from the alarm set point, that is necessary to clear the alarm. Simply stated, the alarm deadband should not be set to zero. This can help to eliminate chattering and fleeting alarms. For example, if the alarm for a pressure transmitter is set at 70 psi and the measured value goes above 70 PSIG, an alarm is generated. Without a deadband, if the value fluctuates around 70 PSIG, an alarm is generated every time the value crosses 70 psig. With a deadband setting equal to say 2 PSIG, the alarm once triggered will not be generated again, till the value drops below 68 PSIG. However, care must be taken when setting the deadband.  If the deadband is set large, this can lead to stale alarms. 

alarm-deadband
Alarm Deadband Settings

ON-OFF DELAYS – Another way to eliminate chattering alarms is to set up an on-delay or off-delay timer on the alarm. With an on-delay timer, the alarm is not annunciated if it does not stay above the set point for a certain time, defined by the on-delay timer. The on-delay timer can be used to fix both chattering and fleeting alarms. The values set should be selected with an understanding of the process dynamics, since an on-delay timer is delaying the time the operator has to respond to an alarm.

Similarly, an off-delay timer sets a duration that the measured process variable has to remain in the normal operating envelope after an alarm is generated before the alarm can be cleared. The off-delay timer is useful for dealing with chattering alarms. 

on-off-time-delay
On-Off Time Delay Settings

The deadband is a more powerful way to deal with nuisance alarms followed by on-delay timers and then off-delay timers. 

alarm-cond
Alarm Conditioning

Alarm Suppression is a useful function for helping to ensure that operators are not presented with alarms unless they are relevant to the current operating mode of the plant. According to the ISA 18.2 standard, all suppressed alarms should be visually indicated in the HMI. However, the visual indication should not include a blinking element and should be distinguishable from the unacknowledged and acknowledged state indications for normal alarms. Most control systems include the option to generate a separate list of suppressed alarms.  

There are three different types of suppression defined in ISA 18.2 – Alarm Shelving, Suppression by Design and Out-of-Service.

Alarm Shelving is a temporary suppression of an alarm, typically initiated by the operator. Alarm shelving enables operators to temporarily remove nuisance alarms until the underlying problem can be addressed. Alarm shelving is different from disabling the alarm. It is manually initiated following a controlled methodology and is under the control of the Operator in the shelved state. The most common application of Alarm Shelving is to manage nuisance alarms from faulty devices.

Most control system vendors offer alarm shelving capabilities in their systems. When an alarm is shelved, it is temporarily moved from the alarm list display to a shelved alarm list display. It stays within this shelved alarm list until it is either cleared by the operator or cleared when the maximum shelving period set for the alarm is reached. The shelved alarms list is recorded and tracked in the alarm system.

State-based suppression is a technique used to suppress alarms based on operating conditions or plant states. Alarm suppression is achieved within the logic of the control system, which evaluates the relevance of the alarm depending on the operating mode of the process area and/or process units. 

State-based alarm suppression can be effective at preventing stale alarms. For example, a low-flow alarm is not useful or relevant when it is generated because a discharge pump has tripped. In this case, the operator’s response and attention should be focused on addressing the pump trip, not on the low flow, which is a consequence of the trip. 

Out-of-service is the alarm state used to manually suppress alarms, when they are removed from service, typically for maintenance. An alarm in the out-of-service state is under the control of maintenance. The control system has the functionality to remove the alarm from service when the associated equipment is out of service. 

What about the lifecycle?

It is possible to perform all the mentioned alarm conditioning techniques to deal with nuisance alarms by undertaking activities that fall within the alarm management lifecycle.

The earlier section of the article stated that the lifecycle can start at the MONITOR AND ASSESSMENT stage. The Engineer can start with this stage of the lifecycle to monitor for nuisance alarms by examining the bad actors in the alarm lists. Let’s say the Engineer  IDENTIFIES a fleeting alarm. The Engineer can then review any RATIONALIZATION or other documentation for the alarm in the Master Alarm Database. The Engineering also discusses the alarm with the maintenance team to ensure there are no control system hardware or field instrumentation issues for this particular instrument loop.  Since the Engineer’s objective is not to decide whether or not to eliminate the alarm, but to solve the nuisance alarm issue, a detailed rationalisation exercise is not required. 

The Engineer will then examine the configuration of the alarm and the historical data for the alarm within the control system as part of the DETAILED DESIGN step. The engineering notices that the deadband has been set to zero. After discussing the alarm with the process and operations teams, the Engineer determines the appropriate deadband value for the alarm. The Engineer will also consult with the  OPERATIONS AND MAINTENANCE (O&M) team, to verify that the changes can be made without any overrides or inhibits. The changes are also recorded within the Master Alarm Database as part of the MANAGEMENT OF CHANGE (MOC). Once the changes and implementation steps are confirmed by O&M teams and approved via the MOC, the change will be IMPLEMENTED online by the systems engineers. The change will then be MONITORED by the Engineer to ensure that the new deadband settings have solved the nuisance alarm issue. 

In this manner, the Engineer is ensuring that the necessary stages of the Alarm Management Lifecycle are being followed while performing the different steps to eliminate the nuisance alarm.

The Benefits

This article has described how the alarm system can be monitored to identify nuisance alarms and how these nuisance alarms can be managed. With these simple measures, the unnecessary load on the operators can be reduced by 50% or more.

The data gathered by monitoring alarms and the benefits to the facility after eliminating nuisance alarms can help build a case for sponsoring an alarm system improvement program. 

The methodologies outlined for resolving nuisance alarms are not a substitute for proper maintenance. Nuisance alarms as a result of faulty hardware or miscalibrated instrumentation should be first resolved by replacing, repairing or recalibrating the faulty equipment. Therefore, regularly reviewing nuisance alarms can often reveal maintenance needs such as instrumentation which is faulty or incorrectly specified for the service.

It is important to understand that managing nuisance alarms does not exempt the plant from maintaining an alarm management philosophy and performing alarm rationalisation to ensure that all alarms present in the system fit the definition of the standard for an alarm. In addition, some stale and redundant alarms can only be removed from the system after alarm rationalisation is carried out. 

That being said, addressing alarm management issues by identifying the bad actors and investigating whether they are nuisance alarms is a good starting point that can help an operating plant achieve a quick win in reducing the number of irrelevant alarms, thereby improving situational awareness for the operators. 

Further Reading

  1. The explosion and fires at the Texaco Refinery, Milford Haven, July 1994. – HSE Books
  2. Better Alarm Handling – HSE Information Sheet
  3. The Management of Alarm Systems – by M.L.Barnsby and J.Jenkinson – HSE Report, 1998
  4. Inspectors Toolkit: Human Factors in Major Hazards – HSE, October 2005
  5. EEMUA Publication 191 Alarm systems – a guide to design, management and procurement 
  6. The Alarm Management Handbook – Second Edition: A Comprehensive Guide – by Bill Hollifield and Eddie Habibi
  7. Get a life(cycle)! Connecting Alarm Management and Safety Instrumented Systems – by Todd Stauffer, Nicholas P. Sands and Donald G. Dunn 
  8. Understanding and Applying the ANSI / ISA 18.2 Alarm Management Standard – by Bill Hollifield
  9. Chemical Safety Board Investigations
  10. BP Texas City Refinery Incident
  11. BP Grangemouth Petrochemical Complex Incident
  12. Buncefield Oil Storage Depot Incident