Sandeep R. Kondaveeti, PhD, P.Eng., SM-ISA, Senior Reliability Engineer and Automation Team Lead, CNOOC Petroleum North America, discuss transforming alarm data and operator actions into operational intelligence for oil and gas facilities.
Oil and gas facilities operate some of the most complex and tightly constrained industrial processes in the world. From upstream production and midstream compression to downstream processing, operators must manage high pressures, high temperatures, flammable hydrocarbons, and strict environmental and safety regulations. Control rooms are expected to maintain stable operation while responding quickly to abnormal situations such as equipment failures, process upsets, and utility disturbances.
Yet in many oil and gas control rooms, alarm management and Advanced Process Control (APC) are still treated as separate disciplines. Alarm management is often approached as a compliance activity that is focused on meeting standards such as ISA-18.2, while APC is implemented primarily to improve throughput, energy efficiency, or product quality during normal operation.
Where alarm management and control strategy fall out of sync
In practice, these two systems are inseparable. When a process upset occurs, alarms are what first alert operators to trouble. Operators then translate those alarms into manual actions – changing setpoints, switching controllers to manual, or adjusting flows and pressures. These actions, in turn, directly affect how APC behaves. Poorly designed alarms increase operator workload, delay corrective action, and often force APC systems into fallback or constraint-handling modes. The result is more alarms, lost production, higher risk, and reduced confidence in automation.
For critical energy infrastructure, this disconnect represents a missed opportunity. Alarm systems and operator actions together contain a rich source of operational intelligence that can be used to strengthen control strategies and improve safety and reliability.
Why alarm metrics alone are not enough
Most alarm programmes focus on measurable quantities such as average alarm rates, standing alarms, chattering alarms, and priority distributions. These metrics answer an important question: how many alarms are occurring? But metrics alone do not answer a more operationally relevant question: which alarms actually drive operator decisions during abnormal situations?
In operating plants, experienced operators quickly learn which alarms matter and which can be safely ignored. Over time, this knowledge becomes embedded in human behaviour rather than in the control system. When alarm rationalisation is performed without capturing this behavioural knowledge, organisations improve statistics but often fail to improve actual response quality during upsets.
Operator actions: an untapped source of intelligence
Modern distributed control systems log far more than just alarms. They record manual controller output changes, mode changes, setpoint adjustments, alarm acknowledgements, alarm shelving or suppression, and bypass and inhibit actions.
When these actions are analysed alongside alarm data, consistent response patterns emerge. For example, a high-pressure alarm on a compressor train may consistently be followed by throughput reduction. A temperature deviation alarm in a separation unit may trigger steam valve trimming. Certain alarms may be acknowledged repeatedly with no corrective action at all.
These patterns form a behavioural model of how operators stabilise the process during abnormal situations. Instead of relying on undocumented experience, this operational knowledge becomes observable and measurable.
How poor alarms undermine advanced process control
Alarm systems and APC influence each other more directly than is often recognised. Poorly tuned control loops generate oscillations that trigger nuisance alarms. Nuisance alarms reduce operator trust in the system. Operators respond by switching controllers to manual or overriding APC. APC performance degrades, creating more instability and more alarms.
This creates a reinforcing loop that undermines both safety and optimisation objectives. Breaking this cycle requires treating alarm systems, operator behaviour, and APC strategies as parts of a single operational ecosystem rather than independent engineering tasks.
A practical framework for oil and gas facilities
Seeking a proactive yet manageable solution to this problem, I developed the following four-stage framework that can be easily implemented and used to convert alarm and operator action data into reusable operational intelligence for oil and gas plants.
Stage 1: characterise alarm events
In the first stage, alarm data is grouped into meaningful structures such as floods versus isolated alarms, repeating vs transient alarms, and alarms associated with specific equipment or process units. Although individual alarms often appear unrelated in raw event logs, the same disturbance typically produces the same sequence of alarms each time it occurs.
By examining these sequences rather than isolated alarm points, it becomes possible to distinguish initiating alarms from cascading symptom alarms. This step reveals which alarms are indicators of an underlying problem and which are merely consequences of it. The key outcome of this stage is recognition that alarm floods are not random – they are repeatable signatures of specific operating disturbances.
Stage 2: map operator responses
Operator actions are aligned with alarm timelines, including time to first intervention, type of corrective action, and whether the alarm clears after action. This makes it possible to see which alarms consistently drive corrective behaviour and which are routinely acknowledged without meaningful response.
Over multiple events, patterns emerge showing how operators interpret alarm information and decide when to intervene. Some alarms reliably trigger action, while others contribute primarily to background noise. Thus, this stage exposes how human judgment filters alarm information into operational decisions.
Stage 3: extract operational rules
This step converts messy event logs into structured guidance that reflects how the plant is actually stabilised during abnormal situations. Repeated alarm–action relationships are translated into operational rules such as “when this alarm pattern occurs, operators reduce throughput” or “these alarms indicate the same underlying disturbance.” Instead of remaining informal or tribal knowledge, successful responses become explicit and transferable. At this stage, experience is transformed into reusable logic that can guide alarm priorities, alarm limits, and control strategy design. The result is not a theoretical model of plant behaviour, but a data-driven representation of how disturbances are managed in practice.
Stage 4: enhance control strategies
Lastly, once operational rules are defined, they can be embedded into control strategies so that known disturbance patterns are addressed automatically. Alarm-derived intelligence influences how the control system adjusts throughput, constraints, or operating modes, reducing reliance on manual intervention and stabilising the process more quickly.
In model predictive control systems, this intelligence can be implemented as soft constraints, disturbance feed forward terms, or mode-switching logic tied to recognised alarm patterns. Control behaviour evolves from purely model-driven action to action informed by historical operator success.
Scaling across fields and facilities
This approach is especially valuable for oil and gas organisations operating fleets of similar assets such as well pads, compressor stations, processing trains, and terminals, where alarm patterns and operator responses can be compared across units to identify best practices. Once encoded into control strategies, these practices become portable rather than local knowledge.
In practical applications, I have found that this approach typically leads to significant reduction in standing and nuisance alarms, faster operator response during disturbances, lower reliance on manual control, and more stable APC performance.
Reducing cognitive load while strengthening process safety
From an operator’s perspective, integrating alarm and action intelligence produces fewer but more meaningful alarms, faster diagnosis of abnormal situations, and reduced cognitive load. Operators are no longer required to mentally filter large volumes of alarm noise to determine what matters. Instead, alarms increasingly reflect conditions that require intervention, and control behaviour aligns more closely with how disturbances are actually mitigated in practice.
This shift changes the nature of work in the control room. Operators move from reactive firefighting to supervisory control, spending less time suppressing symptoms and more time managing process objectives. In oil and gas facilities, where disturbances can propagate quickly across units and margins for error are small, this reduction in cognitive burden directly supports process safety. Fewer unnecessary interventions, clearer decision cues, and more stable control responses reduce the likelihood of delayed or incorrect actions during high-risk situations involving hydrocarbons, pressure systems, and thermal equipment.
Compliance and governance implications
Regulators and corporate governance programs increasingly emphasise performance-based evidence of alarm system effectiveness, rather than simple documentation of alarm limits and rationalisation activities. By linking alarm behaviour with operator response and disturbance recovery, facilities can demonstrate that alarms exist, and that they meaningfully support safe and stable operations.
This approach strengthens key lifecycle activities such as ongoing performance monitoring, management of change, and periodic alarm assessments. Changes to process equipment, control strategies, or operating conditions can be evaluated based on how they alter alarm patterns and operator behaviour, rather than on alarm counts alone. Over time, alarm metrics evolve from static indicators of compliance into dynamic indicators of operational risk, providing both internal governance teams and external auditors with clearer evidence that alarm systems are exceeding the mere satisfaction of procedural requirements by better contributing to abnormal situation management.
Toward smarter control rooms for critical energy infrastructure
The long-term vision is a control room where alarm systems adapt to evolving process conditions, control strategies incorporate operator expertise, and operator actions refine system intelligence over time. In such an environment, alarms become guidance rather than noise, operators become supervisors rather than firefighters, and APC becomes adaptive rather than static.
In oil and gas facilities, alarm systems, operator behaviour, and advanced control strategies are tightly coupled elements of safe and reliable operation. By analysing alarms as well as how operators respond to them, organisations can extract operational intelligence that improves alarm quality, strengthens APC performance, and enhances process safety. For critical energy infrastructure, the next evolution of alarm management goes beyond reducing alarm counts. Smarter alarms are those informed by real operator behaviour, real disturbances, and the realities of how oil and gas plants are run when conditions are no longer normal.
About the author
Sandeep R. Kondaveeti is a Senior Reliability Engineer and Automation Team Lead with CNOOC Petroleum North America, specialising in industrial automation and control systems for critical energy infrastructure. With more than 15 years of experience in oil and gas, power generation, and process automation, his work focuses on advanced process control, industrial alarm management, and safety-critical control systems, with an emphasis on sustaining operational performance beyond project execution. He has led alarm management and compliance programs delivering measurable improvements in safety, reliability, and efficiency, and serves as a Responsible Engineer for critical infrastructure systems. Sandeep holds a Ph.D. in Chemical and Materials Engineering with a focus on computer process control.