Root Cause Analysis (RCA) applied to engineering: methodology, evidence, 5 Whys, Ishikawa, barriers, hypotheses, corrective actions, and effectiveness verification.
Check it out!
Root Cause Analysis — RCA — is a structured process for investigating why a failure, incident, nonconformity, or loss of performance occurred and which conditions allowed it to materialize. The objective is not to find a convenient explanation or assign blame, but to build an evidence-based causal chain and identify actions capable of reducing the probability of recurrence.
In reliability engineering, RCA is especially relevant when simply restoring function does not solve the problem. Replacing a burned-out component, restarting a controller, or retightening a connection may recover the system, but it does not explain why the event occurred or whether the same mechanism will remain present. RCA begins precisely where corrective maintenance ends: after service is restored, the mechanism, contributing conditions, failed barriers, and associated systemic causes are investigated.
RCA is not synonymous with Ishikawa, 5 Whys, or fault tree analysis. These are techniques that can support an investigation. RCA is the complete process: define the event, preserve evidence, reconstruct the sequence, formulate hypotheses, test causality, identify physical, human, and organizational causes, propose actions, and verify whether they effectively reduce the risk of recurrence.
What is Root Cause Analysis?
IEC 62740:2015 describes RCA as a retrospective process for analyzing events that have occurred, applicable to failures, incidents, nonconformities, and other relevant events. The standard emphasizes that causes may be related to design, processes, organizational factors, human aspects, and external events. This is important because technical failures rarely belong only to the component that stopped.
In an electrical installation, for example, repeated circuit-breaker tripping may be caused by overload, inadequate selectivity, incorrect settings, degradation, harmonics, a downstream fault, environmental conditions, a design error, or an operational intervention. RCA must separate symptom, mechanism, and cause.
A practical definition is: a root cause is a causal condition that, when properly addressed, materially reduces the probability of recurrence of the event or equivalent events.
RCA does not look for someone to blame; it looks for a causal chain that can be changed. If the conclusion does not change recurrence or risk, the investigation probably stopped too early.
This avoids a common error: calling any fact at the beginning of the narrative a root cause. "The relay failed" may only describe the last visible link in a causal chain.
Events, failures, and problems must be defined precisely
A poor investigation often begins with a poorly defined event. Expressions such as "the system went down," "the pump stopped," or "the panel failed" are insufficient.
Whenever possible, the event should indicate:
- the function that was lost or degraded;
- the equipment or system affected;
- the instant or time window;
- the operating condition;
- the observed consequence;
- duration;
- the state before and after the event;
- relevant recent changes.
A better example: total loss of pumping function in circuit X at 2:32 p.m., during normal-load operation, after simultaneous operation of both motor protection devices, resulting in 47 minutes of process unavailability.
The more precise the definition, the more objective the search for evidence will be.
RCA begins with evidence preservation
The natural impulse after a failure is to restore operation quickly. That is correct from an operational standpoint, but it may destroy data needed for the investigation.
Before disassembling, resetting, replacing, or changing settings, the team should assess which evidence must be preserved. Depending on the system, this may include:
- logs and alarms;
- oscillography and protection records;
- process trends;
- photographs and videos;
- the state of relays, contactors, and protection devices;
- valve and switch positions;
- software settings and versions;
- physical samples;
- temperature, vibration, or current measurements;
- previous maintenance work orders;
- operator statements;
- records of recent interventions.
The time sequence is especially important. In digital systems, events may occur in milliseconds; in mechanical failures, degradation may develop over months.
Symptom, failure mechanism, and cause are not the same thing
Separating these three levels improves RCA quality.
| Level | Question | Example |
| Symptom | What was observed? | motor stopped |
| Mode/mechanism | How was the function lost? | bearing seized after temperature increased |
| Cause | Why did the mechanism develop? | inadequate lubrication due to an incorrect procedure and incompatible interval |
Replacing the bearing addresses the damage. Correcting only the interval may address part of the mechanism. If the procedure does not define quantity, specification, method, and evidence, recurrence may still remain possible.
Immediate, contributing, and systemic causes
A robust RCA often finds multiple causal levels.
Immediate cause is the mechanism closest to the event: short circuit, loss of lubrication, unintended command, rupture, overheating.
Contributing conditions increase probability or consequence: severe environment, poor access, ineffective alarm, outdated documentation, lack of spare parts.
Systemic causes are related to design, process, management, competence, governance, or control decisions that allowed the condition to exist and persist.
This classification avoids the trap of ending the investigation at "human error." An operational error may be part of the chain, but RCA should investigate why the system allowed a single error to produce the consequence, whether barriers existed, whether the interface was clear, and whether the procedure was executable.
How to structure an RCA step by step
A pragmatic process can follow nine steps:
1. Define the event and impact. Record what happened, when, under what condition, and which function was lost. 2. Preserve and collect evidence. Ensure traceability of physical, digital, and documentary data. 3. Reconstruct the timeline. Order facts and separate observation from interpretation. 4. Identify failure modes and mechanisms. Understand technically how the event materialized. 5. Formulate causal hypotheses. Create plausible explanations without prematurely selecting a favorite. 6. Test hypotheses against evidence. Look for confirmation as well as evidence that can refute them. 7. Identify causes and contributing factors. Cover physical, human, organizational, and external dimensions. 8. Define corrective and preventive actions. Prioritize actions that effectively modify the causal chain. 9. Verify effectiveness. Track whether the actions reduced recurrence, exposure, and risk.
The process should be proportional to the consequence of the event. A three-week RCA for a trivial failure may be wasteful; a superficial investigation after a critical failure may leave the organization exposed.
Build the timeline before the causal narrative
Reconstructing the factual sequence before discussing causes is a useful discipline. The timeline may combine:
- previous normal condition;
- recent changes;
- first detectable deviation;
- alarms;
- automatic actions;
- human decisions;
- functional failure;
- emergency response;
- restoration.
This method reduces hindsight bias. Once we know the outcome, it is easy to interpret each previous event as "obvious." The timeline preserves what was actually observable at each point in time.
5 Whys: useful for digging deeper, insufficient on its own
The 5 Whys technique encourages the team not to stop at the first explanation. Its value lies in the discipline of exploring causality in greater depth.
Simplified example:
1. Why did the pump stop? — the motor tripped on overtemperature. 2. Why did it overheat? — ventilation was obstructed. 3. Why was it obstructed? — particulate material accumulated. 4. Why was the buildup not identified? — the inspection did not include that point. 5. Why did the inspection not include it? — the plan was created from a generic manual without considering the actual environment.
The problem is turning "five" into a rule. Some chains require three levels; others require fifteen. In addition, complex events have parallel and interdependent causes that a linear chain does not represent well.
For the specific application of the technique, including use criteria, examples, limitations, and integration with corrective action, see 5 Whys in Engineering: how to investigate root cause without stopping at the symptom.
Ishikawa Diagram as an exploration tool
The Ishikawa Diagram organizes hypotheses by category and helps avoid a narrow investigation. It may consider machine, method, manpower, material, measurement, environment, or categories adapted to the system.
It is particularly useful at the beginning, when the team needs to broaden the hypothesis space. However, an item placed on the diagram does not become a cause merely because it was remembered. Each hypothesis must be confronted with evidence.
On A3A’s website, the dedicated article on the Ishikawa Diagram remains responsible for teaching the tool. On this page, Ishikawa appears as one technique within the broader RCA process.
Cause trees and causal logic
When an event results from a combination of several factors, cause trees can represent AND/OR relationships and dependencies among conditions.
Consider a power-supply failure in which the load is lost only when:
- the primary source is unavailable; and
- automatic transfer does not operate; and
- manual bypass cannot be executed within the required time.
The cause of the event cannot be reduced only to the failure of the primary source. The system was designed precisely to tolerate that condition. The investigation must explain why the defense barriers also failed.
FTA and RCA are not the same thing
FTA — Fault Tree Analysis — is a deductive technique that starts from a top event and decomposes combinations of failures capable of producing it. It may be used prospectively or to structure hypotheses.
RCA analyzes an event that actually occurred. It uses evidence from the case and may incorporate FTA, Ishikawa, 5 Whys, barrier analysis, change analysis, and other techniques.
The choice depends on complexity, criticality, and the nature of the evidence.
Barrier analysis
A powerful way to investigate incidents is to ask which barriers should have prevented the event or reduced its consequence.
Barriers may be:
- physical: protection, interlocking, containment;
- automatic: logic, alarm, shutdown;
- procedural: checklist, permit, inspection;
- human: independent review, double-checking;
- organizational: management of change, competence, technical approval.
For each barrier, the team assesses whether it existed, whether it was adequate, whether it was available, and whether it performed as expected.
This approach shifts the investigation from "who made the mistake" to "how did the system allow the event to evolve."
A barrier that existed in the procedure but could not be executed in the field is not an effective barrier. RCA must assess the existence, adequacy, availability, and actual performance of defenses.
Change analysis
Events often arise after an explicit or silent change: a new supplier, process adjustment, software update, material replacement, load change, team change, procedure revision, or layout modification.
Comparing before vs. after helps identify relevant variables. However, temporal correlation does not prove causation. The change must be technically connected to the observed mechanism.
How to test a causal hypothesis
A good hypothesis should explain the event and be consistent with the evidence.
Useful questions include:
- does the hypothesis explain the time sequence?
- does it explain the observed physical damage?
- is it consistent with measurements and logs?
- is there a plausible technical mechanism?
- have equivalent events occurred before?
- would the system operate normally if this condition were absent?
- is there evidence that contradicts the hypothesis?
The last question is critical. A reliable investigation tries to refute its own hypothesis, not merely accumulate elements that confirm it.
Technical example: recurring power-supply failure
Consider an industrial power supply that fails three times in six months. The corrective action each time was to replace the module.
The RCA identifies the following facts:
- all failures occurred in the same panel;
- the modules show similar damage in the input stage;
- measurements show transients above the expected condition;
- the installed surge protective device is at end of life and has no monitored indication;
- the design does not provide adequate coordination among protection devices;
- preventive inspection checks only physical presence, not functional status.
The "root cause" is not simply a "defective power supply." The mechanism is associated with electrical exposure and inadequate protection/monitoring barriers.
Possible actions include reviewing protection coordination, replacing the degraded device, improving diagnostics, updating inspection criteria, and checking other installations with the same architecture.
This last point is important: a good RCA does not address only the asset that failed; it looks for replicated systemic risk.
Corrective action must address the causal chain
Actions can be classified by the level at which they operate.
| Action type | Example | Typical strength |
| restore | replace component | low against recurrence |
| detect | create alarm or inspection | medium |
| reduce exposure | change procedure or frequency | medium |
| eliminate mechanism | correct design, process, or material | high |
| create independent barrier | interlocking, protection, segregation | high |
This does not mean every action must involve major engineering work. In many cases, the best solution is simple. The criterion is whether it defensibly changes the probability or consequence of the event.
Traceable action plan
Each action should record:
- the cause or factor it is intended to address;
- responsible person;
- deadline;
- implementation evidence;
- residual risk;
- criterion for verifying effectiveness.
"Train the team" is insufficient unless it is clear which behavior must change, why training is the appropriate barrier, and how the result will be verified.
Effectiveness verification
Closing the administrative action does not close the RCA. The organization must verify whether the condition has actually been controlled.
Indicators may include:
- recurrence of the failure mode;
- failure rate after intervention;
- availability;
- reduction in alarms/precursor events;
- inspection results;
- compliance with new parameters;
- elimination of the mechanism in equivalent assets.
For rare events, waiting for a new failure may not be acceptable. In that case, effectiveness may be demonstrated through testing, inspection, calculation, design review, or barrier verification.
When to open a formal RCA
Not every failure deserves the same level of effort. Trigger criteria may include:
- relevant accident or near miss;
- loss of a critical function;
- failure with environmental or regulatory impact;
- unavailability above a defined limit;
- recurrence;
- high cost;
- failure of a critical barrier;
- unexpected event in a redundant system;
- risk of repetition in similar assets.
Criticality helps calibrate the depth of analysis and the team required.
Who should participate
A multidisciplinary RCA tends to be more robust. Depending on the event, participants may include operations, maintenance, engineering, automation, safety, quality, the manufacturer, and external specialists.
The team should combine system knowledge with enough independence to challenge assumptions. When the investigation is conducted only by those who designed or implemented the solution, confirmation bias may arise.
Common RCA mistakes
Some patterns drastically reduce investigation quality:
- selecting the cause before collecting data;
- confusing correlation with causation;
- stopping at "human error";
- using interviews alone without technical evidence;
- applying 5 Whys mechanically;
- listing dozens of causes without prioritization;
- proposing training for every problem;
- closing the analysis after replacing the part;
- failing to verify whether the same condition exists in equivalent assets;
- failing to measure action effectiveness.
RCA, FMEA, and FMECA complement one another
FMEA and FMECA are primarily structured to anticipate failure modes and effects or to analyze a system systematically. RCA starts from an event that has occurred and reconstructs causality.
A well-conducted RCA can feed FMEA/FMECA with new modes, causes, controls, and evidence. Likewise, an existing FMEA can accelerate the investigation by providing hypotheses and functional relationships that have already been analyzed.
This cycle turns an incident into engineering knowledge.
RCA and FRACAS
FRACAS — Failure Reporting, Analysis and Corrective Action System — extends RCA into a continuous process of recording, analysis, action, and closure. While RCA may be a specific investigation, FRACAS repeatedly organizes events throughout the life of the system.
Maturity appears when causes and actions no longer remain isolated in reports and instead feed databases, design standards, maintenance, training, and asset decisions.
When RCA adds the most value
RCA is especially valuable when there is recurrence, high impact, uncertainty about the mechanism, multiple barriers involved, or risk of replication. It is also useful in commissioning, assisted operation, maintenance, failures of critical systems, and analysis of below-expected performance.
The deliverable should support decision-making: well-defined event → evidence → causal chain → causes → actions → effectiveness verification. Without this chain, the investigation tends to become a retrospective narrative rather than an engineering tool.
Closing the action does not mean closing the cause. RCA completes the cycle only when there is implementation evidence and a technical criterion for verifying effectiveness.
Technical references
[1] IEC. IEC 62740:2015 — Root cause analysis (RCA). Geneva: IEC, 2015.
[2] IEC. IEC 60812:2018 — Failure modes and effects analysis (FMEA and FMECA). Geneva: IEC, 2018.
[3] ISO. ISO 31000:2018 — Risk management — Guidelines. Geneva: ISO, 2018.
Frequently asked questions
It is a structured process for investigating events that have occurred, identifying evidence-based causes and contributing factors, and defining actions capable of reducing the probability of recurrence.
No. 5 Whys is a technique that can support the investigation. RCA is the complete process of event definition, evidence collection, causal analysis, actions, and effectiveness verification.
No. Ishikawa helps organize cause hypotheses; RCA requires testing those hypotheses against evidence and building a defensible causal chain.
When the event has high consequences, recurrence, critical-barrier failure, regulatory impact, major unavailability, high cost, or risk of repetition in similar assets.
It can be part of the causal chain, but it is usually insufficient to end the analysis at that level. The investigation should assess design conditions, interface, procedures, training, barriers, and organizational factors that allowed the error to produce the consequence.
Effectiveness should be verified through indicators, tests, inspections, calculations, or evidence demonstrating reduction of the mechanism, recurrence, or associated risk.
Complementary technical materials
Related solutions
- Field Applications, Inspection, and Technical Data Collection
- Technical Knowledge Management and Lessons Learned
- Requirements, Evidence, and Acceptance Criteria Management
Related engineering services
- Reliability and Availability Engineering
- Maintenance Engineering
- Recommissioning of Systems and Facilities
Related technical content
Guides, frameworks, and references
