FMEA applied to engineering: functions, failure modes, effects, causes, controls, severity, occurrence, detection, RPN, Action Priority and applications.
Check it out!
FMEA — Failure Modes and Effects Analysis — is a systematic method for identifying how an item, system, design or process can fail, which effects those failures may produce, which causes are associated with them and which actions should be prioritized to reduce risk before the problem materializes.
The strength of the method lies less in the spreadsheet and more in the reasoning process. A well-conducted FMEA forces the team to clarify functions, interfaces, requirements, potential failures, causal mechanisms, existing controls and evidence. This makes it possible to transform dispersed expert knowledge into a structured and traceable analysis.
Although it is widely known in the automotive industry, FMEA is not limited to that sector. IEC 60812:2018 establishes a generic approach applicable to hardware, software, processes, human actions and interfaces. The AIAG & VDA handbook provides a specific harmonized methodology for automotive applications of Design FMEA, Process FMEA and FMEA-MSR.
What Is FMEA
FMEA starts with a simple question: in what ways can this function fail to be fulfilled?
From there, the team identifies failure modes, effects and causes and evaluates which situations require treatment. The method can be applied before implementation, during design reviews, in production processes, during modifications, in maintenance engineering or in the analysis of existing systems.
According to IEC 60812:2018, FMEA provides a systematic method for identifying failure modes and their local and overall effects, and may also incorporate their causes. Failure modes can be prioritized to support treatment decisions.
A typical analysis connects six elements:
- function or requirement;
- failure mode;
- failure effect;
- failure cause or mechanism;
- existing controls;
- action required to reduce or control risk.
When these elements are treated only as form columns, FMEA tends to become bureaucratic. When they are treated as an engineering logic chain, the method helps reveal risks that have not yet appeared in the field.
A useful FMEA starts with the function and ends with a verifiable action. The spreadsheet is only the record; the value lies in the logic connecting requirement, failure mode, effect, cause, control and decision.
Function Comes Before Failure
One of the most common mistakes is to begin an FMEA by listing known defects. The analysis should begin with the function.
Consider a pumping system. A “burned-out motor” is a possible physical failure, but it does not necessarily describe the system’s functional failure. Depending on the scope, the function may be to maintain a specified flow, pressure or availability. The system can fail to fulfill that function even when the motor has not burned out — for example, because of cavitation, blockage, a control error, loss of power, an incorrect sensor or improper configuration.
Defining the function makes it possible to identify failure modes more completely and reduces the risk of limiting the analysis to known history.
What Is a Failure Mode
A failure mode describes how the function ceases to be fulfilled or begins to be fulfilled inadequately.
Examples may include:
- fail to start when required;
- stop during operation;
- operate below required capacity;
- provide incorrect output;
- operate outside the specified timing;
- remain energized when it should shut down;
- leak;
- lose communication;
- produce an incorrect measurement;
- fail intermittently.
The level of detail should be compatible with the decision. Modes that are too generic do not guide actions; modes that are excessively fragmented make the analysis impractical.
Effect, Failure Mode and Cause Are Not the Same Thing
The quality of an FMEA depends on separating these concepts.
Effect is what happens as a consequence of the failure. It may be local, occur at subsystem level or appear at the user/process level.
Failure mode is the way in which the function is lost or degraded.
Cause is what initiates or explains the failure mode within the adopted level of analysis.
A simplified example:
| Element | Example |
| Function | supply a critical load continuously |
| Failure mode | power output unavailable |
| Local effect | load without power |
| System effect | interruption of a critical process |
| Potential cause | component, control, connection or upstream supply failure |
| Control | redundancy, protection, monitoring, periodic testing |
The same occurrence may appear as an effect at one level and as a failure mode at another. The analysis boundary therefore needs to be defined before the FMEA is populated.
Types of FMEA
The fundamental structure is similar, but the object changes according to the application.
Design FMEA — DFMEA
DFMEA evaluates risks associated with product or system design. The focus is on functions, requirements, architecture, interfaces and engineering choices.
It can support decisions about:
- redundancy;
- fault tolerance;
- materials and components;
- protection;
- interfaces;
- capacity;
- diagnostics;
- maintainability;
- test criteria;
- environmental conditions.
The earlier it is applied, the greater the opportunity to eliminate risks through design changes rather than adding later controls to compensate for them.
Process FMEA — PFMEA
PFMEA analyzes how a process may generate inadequate results. It is widely used in manufacturing, assembly, installation, execution and repetitive activities.
The focus may include operation sequence, parameters, resources, tools, human error, inspection, measurement and process controls.
System FMEA
In complex systems, the analysis considers functions and interfaces among subsystems. Shared dependencies, common-cause failures and propagation of effects become especially important.
FMEA Applied to Operations and Maintenance
FMEA can also support the analysis of operating assets, especially when combined with criticality, failure history, condition and maintenance strategies. In this context, it helps identify which modes require prevention, monitoring, detection, contingency or design changes.
FMEA vs. FMECA: What Is the Difference?
FMECA — Failure Modes, Effects and Criticality Analysis — adds a formal criticality assessment to failure modes. IEC 60812:2018 addresses FMEA and FMECA within the same reference framework and allows different prioritization methods.
For the reliability cluster, the distinction is important: FMEA identifies and structures failure risks; FMECA deepens their criticality using defined criteria. FMECA deserves a separate analysis because it may involve criticality matrices, severity and other quantitative or semi-quantitative parameters.
How to Perform an FMEA Step by Step
There is no single universal procedure for every sector, but a robust FMEA follows a logic that should not be reversed: scope → function → functional failure → failure mode → effect → cause/mechanism → controls → assessment → action → verification. When a team starts from a list of components that “can break,” it tends to produce a long spreadsheet that is poorly connected to the system’s actual performance.
ABNT NBR 5462 helps separate concepts that are often mixed together. Failure is the termination of an item’s ability to perform the required function; failure cause describes the circumstances that lead to the event; failure mechanism describes the physical, chemical or other processes that produce it. This distinction is fundamental because effects, causes and mechanisms require different treatments.
| Element | Engineering question | Example |
|---|---|---|
| Function | What must the item do and at what level? | Maintain minimum flow of 10,000 m³/h |
| Functional failure | How can the function cease to be fulfilled? | Flow below minimum |
| Failure mode | Which state or event produces the functional failure? | Fan does not rotate |
| Effect | What happens when the mode occurs? | Room temperature rises |
| Cause/mechanism | Why can the mode occur? | Bearing seized due to loss of lubrication |
| Control | How can the consequence be prevented, detected or limited? | Vibration, temperature and redundancy |
| Action | What should change to reduce risk? | Monitoring, redesign or policy revision |
This sequence improves analysis quality because it makes each line verifiable. If the team cannot state the function and requirement, it still lacks a basis for claiming that a failure occurred. If the cause is described only as “wear,” the mechanism has probably not been explored deeply enough. If the control does not act on the cause or reduce the effect, it should not receive credit in prioritization.
Granularity also needs to be proportional to the decision. An architectural FMEA works with functions and subsystems; a DFMEA may reach components and interfaces; an analysis applied to maintenance should deepen the modes that actually drive tasks, inspections and policies. Detailing every bolt without decision value increases effort without increasing quality.
Define Objective and Scope
Before the analysis, determine what will be studied, the life-cycle phase, boundaries, interfaces, level of decomposition, operating conditions and the decision the FMEA is intended to support.
An FMEA used to review a detailed design is different from one used to define the maintenance strategy for an existing plant.
Structure the System or Process
Decompose the object into coherent levels: system, subsystem, equipment, function, process step or another suitable structure.
Structural analysis reduces gaps and prevents relevant components or interfaces from being left outside the scope.
Identify Functions and Requirements
For each element, record what must be achieved and under which criteria. Whenever possible, use verifiable requirements: capacity, time, operating range, availability, accuracy, protection or another technical parameter.
Identify Failure Modes
Ask how the function can be lost, reduced, exceeded, delayed, intermittent or performed incorrectly.
Identify Effects
Assess what happens locally and how the effect propagates. It is important to reach the consequence that is relevant to the system, process, user or business.
Identify Causes and Mechanisms
Causes should be specific enough to guide actions. “Equipment failure” is rarely a useful cause. Wear, contamination, overtemperature, parameterization error, communication loss, looseness, fatigue, overvoltage, improper installation or an incorrect procedure are more actionable examples when supported by context.
Assess Existing Controls
Controls may act on prevention, detection or mitigation. It is necessary to distinguish design controls, inspection, testing, monitoring, alarms, redundancy, procedures and contingency measures.
Prioritize Risks
Prioritization depends on the reference framework adopted. Traditional methodologies frequently use severity, occurrence and detection and may calculate RPN. The AIAG & VDA approach for the automotive sector introduced Action Priority — AP instead of RPN as the method for prioritizing actions.
This does not mean that every FMEA in every sector should use AP. The organization needs to adopt criteria compatible with its reference framework, contractual requirements and the nature of the risk.
Define Actions and Owners
FMEA creates value only when it produces effective actions. Each action should have an owner, deadline, objective, implementation evidence and a rule for reassessing residual risk.
The AIAG & VDA Seven-Step Approach
For automotive applications, the AIAG & VDA handbook structures FMEA development into seven steps:
- Planning and preparation.
- Structure analysis.
- Function analysis.
- Failure analysis.
- Risk analysis.
- Optimization.
- Results documentation.
The structure reinforces a valuable principle even outside the automotive industry: risk analysis should be preceded by a sound understanding of structure and functions.
However, when the work is not subject to automotive requirements, IEC 60812:2018 provides a more appropriate generic reference for adapting the method to the engineering context.
Severity, Occurrence and Detection
These three criteria are widely associated with FMEA.
Severity assesses the significance of the failure effect.
Occurrence represents the frequency or probability associated with the cause or mode, according to the method used.
Detection considers the ability of controls to identify the cause or mode before the relevant effect occurs or reaches the user, again according to the adopted scale.
The scales should not be improvised. They need objective, consistent definitions appropriate to the sector. The same numeric rating can represent very different risks when organizations use different criteria.
The Problem with Using RPN Alone
RPN — Risk Priority Number — normally results from multiplying severity, occurrence and detection. It is simple for sorting large spreadsheets, but it has an important mathematical limitation: the same product can represent completely different risk profiles.
Consider two modes. Mode A has severity 10, occurrence 2 and detection 3, resulting in RPN 60. Mode B has severity 5, occurrence 6 and detection 2, also resulting in RPN 60. Numerical equality does not imply decision equivalence. The first may represent a rare but unacceptable safety consequence; the second may represent a frequent but reversible operational loss.
Another problem is false linearity. Scales from 1 to 10 are ordinal: severity 10 is not necessarily “twice” severity 5, even though multiplication treats the numbers that way. Small classification changes can also alter ranking without the real risk changing in the same proportion.
The AIAG & VDA evolution adopted Action Priority as the main prioritization mechanism in the automotive context, keeping severity, occurrence and detection visible instead of reducing them to a single product. In other applications, criticality matrices, escalation rules or specific safety and compliance criteria may be more appropriate.
A practical rule is to establish non-compensable triggers. Severe safety, environmental or regulatory consequences may require action regardless of an occurrence estimated as low. Likewise, a mode with high occurrence may justify continuous-improvement action even if the individual effect is moderate.
The general lesson is: the number should support the decision, not replace it. Prioritization needs to preserve visibility of the consequence, confidence in the data, effectiveness of controls and the context in which the function operates.
Prioritizing risk is not the same as sorting a spreadsheet by the largest number. Severity, consequences, controls, uncertainty and context need to remain visible in the technical decision.
FMEA and Reliability Engineering
FMEA is one of the core tools of Reliability Engineering, but it does not solve every reliability problem.
It is predominantly a structured analysis of modes and effects. When a decision requires modeling probability of success, architectural availability, time-dependent behavior, logical combinations of events or life distributions, other methods such as RAM, RBD, FTA and Weibull may be necessary.
The value lies in combining tools according to the engineering question.
FMEA Applied to Maintenance
In maintenance, an FMEA can help review existing tasks and identify which failure modes actually justify intervention.
For each relevant mode, the team can ask:
- is there a detectable degradation mechanism;
- is it technically possible to monitor it;
- is there a P-F interval or other known behavior;
- does a preventive task effectively reduce the probability of failure;
- can the failure be tolerated until it occurs;
- is there a relevant safety, environmental or operational consequence;
- is it better to modify the design than to increase maintenance.
This logic provides a foundation for methods such as RCM. The Maintenance Engineering service can use failure and criticality analyses to structure plans that are more proportional to risk.
FMEA in Engineering Design
During design, FMEA can be incorporated into Design Reviews and technical gates. This is particularly useful for critical and multidisciplinary systems.
A review can select critical functions and verify:
- total or partial loss of function;
- interface failures;
- power-supply or utility failures;
- communication failures;
- fail-safe condition;
- common causes;
- isolation capability;
- maintenance accessibility;
- alarms and diagnostics;
- recovery after failure;
- tests required to demonstrate the controls.
The output may feed requirements, drawings, specifications, automation logic, test plans, spare-parts strategies and operating procedures.
FMEA and Risk Management
FMEA does not replace a complete risk-management process. It is a specific method for risks associated with failure modes.
Contractual, financial, regulatory, schedule, information-security or market risks may require other techniques. The article on Risk Management in Engineering Projects addresses the broader governance of project risks.
When FMEA is used within that system, critical failure modes can be escalated to corporate or project registers while maintaining the connection between technical risk and management decisions.
Common FMEA Mistakes
An FMEA loses value when it is produced only to satisfy a documentation requirement. Some recurring signs are:
- generic or missing functions;
- copying failure modes from another piece of equipment without reviewing the context;
- confusing effect, mode and cause;
- using the same cause for every mode;
- assigning severity, occurrence and detection without defined criteria;
- prioritizing only by RPN;
- recording controls that do not exist or are not tested;
- creating actions without an owner and evidence;
- failing to update the analysis after a design change;
- producing the FMEA individually rather than with a multidisciplinary team;
- ignoring interfaces and common-cause failures.
FMEA should remain a living document while the analyzed object changes.
Who Should Participate
Analysis quality improves when different perspectives are combined. Depending on the scope, the team may include:
- design engineering;
- operations;
- maintenance;
- automation and control;
- safety;
- quality;
- suppliers;
- commissioning;
- process specialists;
- asset-management personnel.
Facilitation also matters. An experienced coordinator maintains the logic connecting function, failure, effect, cause, control and action and prevents the session from becoming an unstructured discussion.
Simplified FMEA Example for a Critical System
Consider a ventilation system for a technical room whose function is to maintain ambient temperature below 27 °C at the design heat load. The function already contains a measurable criterion, so it is possible to distinguish acceptable degradation from functional failure.
| FMEA field | Example |
|---|---|
| Function | Maintain room ≤ 27 °C under design conditions |
| Functional failure | Fail to maintain required temperature |
| Failure mode | Insufficient airflow in the supply circuit |
| Local effect | Reduced heat removal |
| System effect | Overtemperature, alarms and possible equipment shutdown |
| Cause 1 | Fan unavailable due to mechanical failure |
| Cause 2 | Incorrect command or loss of power supply |
| Cause 3 | Filter/duct blockage increasing pressure drop |
| Current controls | Temperature alarm, status indication and periodic inspection |
| Possible actions | Airflow detection, redundancy, maintenance review and functional testing |
Notice that “fan breaks” is not a sufficient line. The fan may be rotating while airflow remains below the requirement because of blockage, incorrect speed, a closed damper or loss of performance. Likewise, loss of one fan may not result in functional failure if another unit automatically takes the load and the remaining capacity meets the requirement.
Assessment of controls also needs to consider when the failure is detected. A high-temperature alarm detects an effect already propagating; an airflow sensor may reveal functional loss earlier; vibration monitoring may act even earlier on a specific fan failure mechanism. These are different layers of prevention and detection and should not receive the same credit without analysis.
If the consequence is loss of critical equipment, the team may conclude that improving detection alone is insufficient and decide on redundancy, electrical segregation or an architectural change. If the consequence is only temporary discomfort in a non-critical area, monitoring and a response procedure may be proportional. FMEA should lead to a decision compatible with the effect, not to an automatic list of actions.
This example shows why FMEA connects requirements engineering, failures, controls, design, commissioning and maintenance. Analysis quality is measured by the ability to explain how a cause leads to loss of function and which action interrupts that chain.
How to Tell Whether an FMEA Is Good
A high-quality FMEA makes it possible to answer quickly:
- which functions are critical;
- which failure modes threaten those functions;
- which effects are most relevant;
- which causes have insufficient controls;
- which actions remain open;
- which risks remain after the actions;
- which design or maintenance decisions were changed by the analysis.
If the team has an extensive spreadsheet but cannot answer these questions, there is probably documentary volume without analytical maturity.
When to Engage Specialized Support
External support can be useful when the system is multidisciplinary, critical, new to the organization or when an independent analysis needs to challenge design assumptions and existing controls.
It can also be relevant for structuring the methodology, defining scales, facilitating workshops, consolidating supplier FMEAs, relating risks to requirements and creating a verifiable action plan.
For existing assets, FMEA can be integrated with criticality, history, inspections and reliability analysis to guide maintenance and modernization priorities.
In critical systems, FMEA should connect with design, maintenance, commissioning and asset management. The method creates value when its actions change requirements, controls, tests and life-cycle decisions.
Technical references
[1] IEC. IEC 60812:2018 — Failure modes and effects analysis (FMEA and FMECA). Geneva: International Electrotechnical Commission, 2018.
[2] AIAG; VDA. AIAG & VDA FMEA Handbook. 1st ed., 2nd printing. Southfield: Automotive Industry Action Group, 2022.
[3] IEC. IEC 60300-1:2024 — Dependability management — Part 1: Managing dependability. Geneva: International Electrotechnical Commission, 2024.
Frequently asked questions
FMEA means Failure Modes and Effects Analysis. The method identifies how functions can fail, their effects, causes, controls and risk-reduction actions.
FMECA adds a formal criticality assessment to failure modes and effects. FMEA structures modes, effects and causes; FMECA deepens prioritization through criticality.
DFMEA analyzes risks in product or system design. PFMEA analyzes risks associated with manufacturing, assembly, installation or execution processes.
They are criteria used in different methodologies to assess the significance of the effect, occurrence of the mode or cause, and the ability of controls to detect the problem according to previously defined scales.
It depends on the adopted reference framework. RPN remains present in many methodologies but has limitations. The automotive AIAG & VDA approach uses Action Priority to prioritize actions.
Whenever relevant changes affect functions, architecture, process, interfaces, causes, controls, requirements or evidence. It should also be reviewed when field failures reveal incomplete assumptions.
Complementary technical materials
Related solutions
- Field Applications, Inspection and Technical Data Collection
- Technical Knowledge and Lessons Learned Management
- Corporate Web Systems for Management, Operations and Integration
Related engineering services
- Reliability and Availability Engineering
- Maintenance Engineering
- Engineering Asset Management
- Systems and Facilities Recommissioning
Related technical content
- Reliability Engineering: Methods and Indicators
- Asset Management: Life Cycle, Value, Risk and Performance
- Risk Management in Engineering Projects
- Design Review in Engineering Projects
- Requirements Management in Engineering
- ISO 55000 and Asset Management
Guides, frameworks and references
