As plants grow larger and more complex, operators are overloaded searching for the source of problems, and while they search, the problem escalates. This whitepaper explains how automated root cause analysis works in a real-time process environment and why manual RCA methods do not scale to complex, multivariate failures.
The executive summary frames the economics, citing published studies on the hourly cost of plant downtime across industrial plants and the daily cost for some oil wells, before noting the additional losses from wasted man-hours and idle energy consumption.
The core of the paper is fault propagation modelling. Once a problem is identified, possible causes are modelled as problems in their own right, linked by arrows showing cause and effect. Working upstream identifies root causes; working downstream predicts impacts. The worked example follows a sudden drop in reformate RON back through fluctuations in reactor delta temperature, chlorides on catalyst, sulfur breakthrough in heavy naphtha feed and desalter water breakthrough, illustrating that one root cause can produce multiple downstream problems and one problem can have many possible root causes.
Automation of this is handled by KnowledgeNet's Root Cause Analysis (KRCA) module. The paper explains that KRCA runs online in real time, whereas fault trees, fishbone diagrams and manual RCA are retrospective, and that KRCA diagrams incorporate multiple root causes and multiple problems rather than one problem at a time. Models are built in a drag, drop and connect graphical editor, so diagrams remain readable by operations, and libraries can be reused across common equipment types.
Supporting mechanisms are covered in detail: events detection and tests returning True, False or Unknown, with event detection typically performed in KRules; a test planner that sequences tests by cost, requesting the cheapest DCS-calculable checks first before manual readings and time-consuming tests; manual tests routed to field operators and confirmed through the GUI; and mitigation versus corrective actions, where mitigation contains effects while the corrective action fixes the root cause, both implemented as messages or as KWorkflow workflows.
Later sections cover KRules design: filters, arithmetic, time series calculations, statistical process control, logic gates, timers and embedded C# custom blocks, plus KPI calculation, continuous improvement through reusable fault models, and a full keyword glossary.