RCM applied to engineering: functions, functional failures, failure modes, consequences, and selection of technically justified maintenance policies.
Check it out!
Reliability-Centered Maintenance (RCM) is a structured process for determining what should be done to ensure that a physical asset continues to perform the required functions in its current operating context. Instead of starting with a list of maintenance tasks, the method starts from functions, identifies how they can be lost, analyzes failure modes and consequences, and selects technically justified policies.
RCM does not mean increasing preventive maintenance or replacing every strategy with predictive monitoring. The result may be a condition-based task, scheduled restoration or replacement, a failure-finding task, a design change, or even a conscious decision not to perform proactive maintenance when the consequences allow it.
This logic makes RCM particularly useful in complex systems, critical assets, and environments where the cost of failures, downtime, safety impacts, or operational loss requires an explicit justification for each maintenance policy.
What Characterizes an RCM Process
SAE JA1011 establishes criteria for evaluating whether a process can be considered RCM. The current edition is SAE JA1011:2024, revised in November 2024. SAE JA1012 serves as a complementary guide explaining how the criteria are applied.
The process needs to address, in a structured manner, questions related to asset functions, functional failures, failure modes, effects, consequences, technically feasible proactive actions, and the treatment of situations in which no suitable proactive task exists.
The essence is to preserve function, not simply the component. Equipment may remain physically operating and still have failed from a functional perspective if it does not deliver the required capacity, accuracy, safety, availability, or performance.
The Seven Classic RCM Questions
The formulation consolidated by the literature and SAE criteria can be organized into seven sequential questions:
- Quais são as funções e padrões de desempenho do ativo no contexto atual?
- De que formas essas funções podem deixar de ser atendidas?
- O que causa cada falha funcional?
- O que acontece quando cada modo de falha ocorre?
- Por que cada falha importa e quais são suas consequências?
- O que pode ser feito para prevenir, detectar ou reduzir a falha?
- O que deve ser feito quando nenhuma tarefa proativa é tecnicamente adequada?
The strength of the method lies in the sequence. Jumping directly to the question “what maintenance should be performed?” without defining function, functional failure, and consequence tends to reproduce historical tasks rather than justify a strategy.
Functions and Performance Standards
A function should express what the user or process expects from the asset. This includes the primary function and, when relevant, secondary functions related to safety, protection, containment, monitoring, control, comfort, environmental integrity, or regulatory requirements.
The required performance standard also needs to be defined. A pump may need to deliver a specified flow; a UPS must support the load within specified conditions; a detection system needs to identify events within coverage and availability requirements.
Without this standard, it becomes difficult to distinguish acceptable degradation from functional failure.
Functional Failure and Failure Mode Are Not the Same Thing
Functional failure describes the inability to fulfill the required function. Failure mode describes the event or mechanism that causes that loss.
A system may have the functional failure “not supply electrical power to the critical load.” Failure modes may include unintended protection opening, internal fault, loss of supply, battery degradation, command error, connection failure, or other specific causes.
This distinction avoids generic plans. Maintenance needs to act on real and technically plausible failure modes, not broad abstractions.
Failure Effects and Consequences
The effect describes what happens when the failure mode occurs: signs, loss of function, secondary damage, recovery time, need for intervention, and possible impacts.
The consequence explains why this matters to the organization. In RCM, consequences are commonly evaluated across dimensions such as safety and environment, operations, non-operational costs, and hidden failures associated with protective functions.
The consequence guides the intensity of the response. The same failure mode may justify different policies in assets with different functions, redundancies, or impacts.
Hidden Failures and Protective Functions
One of the most important aspects of RCM is the treatment of hidden failures. They occur when a protective function fails without immediate evidence during normal operation.
Protection relays, safety devices, alarms, interlocks, emergency systems, and redundant elements may appear available until the moment their function is demanded. Functional tests and failure-finding tasks therefore have a specific role in the strategy.
The absence of observed failures is not sufficient evidence that the protective function is available.
How RCM Selects the Maintenance Policy
When a maintenance plan has accumulated historical tasks without clear justification, RCM makes it possible to rebuild the logic from function, failure, and consequence. The objective is not to increase the number of activities, but to eliminate tasks with no value and reinforce those that actually control risk.
After understanding function, mode, and consequence, the method evaluates whether a technically applicable and effective task exists. This can lead to different policies.
| Policy | When it may be appropriate |
| Condition-based maintenance | Detectable evidence exists before failure with sufficient time to respond |
| Scheduled restoration | Degradation is related to age or use and restoration reduces the probability of failure |
| Scheduled replacement | A technically justifiable useful life exists for the item |
| Failure finding | Hidden or protective functions need to be verified |
| Planned corrective maintenance | The consequence is tolerable and no proactive task adds value |
| Redesign | The consequence is unacceptable and maintenance cannot adequately control the risk |
RCM therefore does not prescribe a technology. It provides the logic for selecting the appropriate technology or policy.
Relationship Between RCM and FMEA/FMECA
FMEA is a powerful tool for identifying failure modes and effects. FMECA adds criticality assessment. RCM uses these analyses within a logic oriented toward preserving functions and selecting policies.
This means that a detailed FMEA can feed an RCM study, but it does not replace the entire process. RCM needs to evaluate consequences and decide whether a task is applicable and effective in the operating context.
The content on FMEA in Engineering and FMECA explores these complementary tools in greater depth.
Relationship Between RCM and Criticality Analysis
Criticality helps prioritize where analytical effort should be applied. Not every asset needs a full RCM study with the same level of detail.
Systems whose failure has major impacts on safety, continuity, production, compliance, or recovery tend to justify greater depth. For simple, low-consequence assets, leaner methods may be sufficient.
Asset Criticality Analysis can therefore serve as a prioritization filter before detailed RCM application.
RCM and Condition-Based Maintenance
CBM is one possible response within the RCM process, not a substitute for the method. If a failure mode produces a detectable signal with useful lead time and a condition-based task effectively reduces risk, it may be selected.
When this does not occur, insisting on monitoring can create false confidence. The correct policy may be scheduled replacement, functional testing, redundancy, or redesign.
This relationship is important because it prevents RCM from becoming synonymous with “sophisticated predictive maintenance.”
The Role of Age and Use
Not all failure modes have a strong relationship with age. In some components, failure probability increases after a known period; in others, failures may occur predominantly at random.
Scheduled replacement or restoration tasks are justified only when there is evidence that failure behavior is related to age or use and that the intervention reduces failure probability or consequence.
Otherwise, replacing components merely because they reached a certain age may consume resources without increasing reliability.
How to Turn an RCM Analysis into an Executable Plan
The study needs to end in implementable policies: task, frequency or trigger, owner, resource, criterion, evidence, and linkage to the asset.
Integration with the Maintenance Plan and Maintenance Planning and Control is what turns the analysis into a controlled operational routine.
Data Required for an RCM Study
If reliable asset registers, history, or documentation consistent with field conditions do not exist, the functional analysis may start from incorrect assumptions. RCM preparation may require surveying and structuring the asset base before the technical workshop.
Study quality depends on the available evidence. Diagrams, design documents, equipment lists, failure histories, work orders, alarms, operating records, manuals, inspection reports, and process data may be relevant.
When the asset register is incomplete or the installation differs from the documentation, the team needs to address that uncertainty before assigning functions and failure modes. In brownfield environments, field surveys and document reviews may be essential preparatory steps.
RCM in Electrical Systems and Critical Infrastructure
Electrical systems, Data Centers, utilities, electronic security, and other critical infrastructure combine physical assets with protective and redundant functions. This makes RCM reasoning particularly useful.
At a substation, for example, it is not enough to think only about circuit-breaker maintenance. Function, protection scheme, auxiliary power, interlocks, selectivity, redundancy, and failure consequences need to be considered. One failure mode may require periodic testing; another may require monitoring; another may require a design review.
This system-level view is consistent with Reliability and Availability Engineering.
When the Result Is Redesign
An important RCM conclusion is recognizing when maintenance cannot adequately control the consequence of a failure.
If no technically applicable proactive task reduces risk to an acceptable level, it may be necessary to change the design, architecture, protection, redundancy, accessibility, specification, or operating strategy.
At this point, RCM stops being only a maintenance tool and connects with design, value engineering, procurement, management of change, and asset modernization.
How to Avoid Bureaucratic RCM
The method loses value when it becomes a large spreadsheet disconnected from operations. This happens when the scope is excessive, participants do not know the system, failure modes are generic, or resulting tasks never reach the maintenance plan and CMMS.
Application should be proportional to criticality, evidence-based, and decision-oriented. Technical facilitation needs to prevent abstract discussions and maintain traceability among function, failure, consequence, and task.
When to Engage Specialized Engineering
Specialized support is useful when the organization needs to review maintenance policies for critical systems, structure functional analysis, integrate different disciplines, validate data, prioritize assets, or convert results into specifications and executable plans.
The work may include Due Diligence, criticality analysis, FMEA/FMECA, RCM facilitation, plan review, reliability, failure analysis, supplier requirements, and implementation support.
When the result requires physical intervention, Engineering can continue through design development, procurement, oversight, testing, and commissioning, maintaining traceability between the identified cause and the implemented solution.
How to Define the Scope of an RCM Study Without Turning Everything into Analysis
RCM does not need to be applied with the same depth to every asset in an organization. The method creates more value when the scope is directed toward systems whose failures have significant consequences or whose current maintenance strategy shows inefficiency, recurring failures, excessive tasks, or low availability. Initial selection may use criticality, history, costs, operational risk, and importance to the process function.
The object of analysis should be defined by function and technical boundary, not only by an asset code. In an electrical system, for example, analyzing only a circuit breaker without considering supply, protection, control, interlocks, and served loads can hide relevant functions. The scope needs to be broad enough to represent the system function and controlled enough to allow detailed analysis.
This boundary definition also avoids duplication. Auxiliary systems, protections, sensors, and power supplies may appear across different equipment, but they need clear ownership in the analysis. The team should know where each function will be addressed, which interfaces are external, and which support failures need to be considered.
Good preparation reduces workshop time and improves decision quality. Diagrams, P&IDs where applicable, single-line diagrams, asset lists, failure histories, maintenance data, manuals, operating parameters, and incident records should be gathered before the sessions. When documentation is inconsistent, the study itself may reveal the need for an asset survey or As-Built update.
Applicability and Effectiveness: The Two Tests an RCM Task Must Pass
A maintenance task should not be selected merely because it appears technically possible. In RCM, it needs to be applicable to the failure mechanism and effective relative to the consequences. Applicability means there is a basis for believing that the task can detect, prevent, or reduce the probability of failure with useful lead time. Effectiveness means that the task outcome justifies its cost, risk, and operational impact for the consequence being analyzed.
A condition-based inspection, for example, is applicable only when there is an observable variable correlated with degradation and a detection window sufficient to act. Periodic replacement is applicable only when the failure mode has a significant relationship with age or use. A detective task makes sense only when there is a hidden function whose unavailability may remain unknown until another failure occurs.
Effectiveness depends on the consequence. A task may technically reduce failure probability but still not be economically rational for a low-criticality, easily replaceable component. For a safety function, on the other hand, effectiveness criteria may be associated with risk reduction and compliance with requirements rather than only direct maintenance cost.
This reasoning helps eliminate traditional tasks that have survived by habit. If an activity has no clearly associated failure mode, does not change probability or consequence, and produces no useful evidence, it should be questioned. RCM is not a method for increasing the amount of maintenance; it is a method for technically justifying what should be done, why, and with what priority.
How to Conduct RCM Workshops with Multidisciplinary Participation
Analysis quality depends on combining Engineering knowledge with operational experience. Designers understand requirements and architecture; maintenance teams know recurring failures and intervention difficulties; operations teams know regimes, alarms, and process effects; safety professionals can assess consequences; suppliers may contribute component-behavior knowledge. The facilitator’s role is to transform these perspectives into functions, functional failures, failure modes, effects, consequences, and traceable decisions.
Excessively large workshops tend to lose focus. The core team should include people who genuinely know the system and have the technical authority to discuss decisions. Additional specialists can be called for specific topics. Preparing information before the meeting prevents using the workshop to search for basic data that could have been gathered beforehand.
The facilitator needs to prevent two common simplifications. The first is converting every failure mode into a periodic preventive task. The second is treating every failure as an isolated equipment event. The process should preserve functional logic and consider systemic consequences, redundancy, hidden failures, interaction with protection, and recovery capability.
Each decision needs to record its justification. If the selected policy is condition-based inspection, it should be clear which parameter will be monitored, at what frequency, and with what intervention criterion. If the decision is planned corrective maintenance, the organization needs to record why the consequence is acceptable and which resources are required for recovery. If the solution is redesign, the finding should generate an Engineering demand and not remain lost in an analysis spreadsheet.
From the RCM Study to the Maintenance Plan and Asset Management
An RCM study creates value only when its decisions reach the management system. Approved tasks need to be converted into executable plans with assets, procedures, frequencies, triggers, resources, competencies, materials, acceptance criteria, and evidence. The result should feed Maintenance Planning and Control or the CMMS and, where applicable, replace legacy tasks that are no longer justified.
Redesign decisions should follow a different workflow. Changes to protection, redundancy, capacity, accessibility, architecture, materials, or technology require requirements development, design, interface analysis, procurement, implementation, and verification. Mixing redesign with routine maintenance orders reduces traceability and may leave systemic problems unresolved.
The study should also be reviewed when new evidence emerges. An unforeseen failure, change in operating regime, design modification, technology replacement, or change in consequence can invalidate the previous strategy. RCM is not a static document; it is a decision logic that should accompany the asset lifecycle.
This connection with asset management is direct. The organization begins to relate functions, risk, performance, maintenance, costs, and renewal decisions. Instead of treating maintenance as a set of isolated routines, RCM provides a framework for justifying policies based on value and consequences for the business.
Final Considerations
RCM is a decision methodology for preserving relevant functions. Its strength lies in requiring the organization to justify each maintenance policy based on functional failures, failure modes, and consequences, avoiding tasks retained only by tradition.
When applied proportionally and integrated with asset management, RCM helps balance preventive actions, condition-based maintenance, planned corrective maintenance, functional testing, and design changes. The result is a maintenance strategy more consistent with risk, performance, availability, and lifecycle cost.
When no maintenance task can reduce a critical consequence to an acceptable level, the technical outcome may be a design change. In this situation, A3A Engenharia can continue from diagnosis through specification, design, procurement, oversight, and commissioning.
Technical references
[1] SAE INTERNATIONAL. SAE JA1011:2024 — Evaluation Criteria for Reliability-Centered Maintenance (RCM) Processes. Available at: https://saemobilus.sae.org/standards/ja1011_202411-evaluation-criteria-reliability-centered-maintenance-rcm-processes.
[2] SAE INTERNATIONAL. SAE JA1012:2011 — A Guide to the Reliability-Centered Maintenance (RCM) Standard. Available at: https://saemobilus.sae.org/standards/ja1012_201108-a-guide-reliability-centered-maintenance-rcm-standard.
[3] NATIONAL AERONAUTICS AND SPACE ADMINISTRATION. Reliability-Centered Maintenance Guide for Facilities and Collateral Equipment. Available at: https://www.nasa.gov/wp-content/uploads/2023/06/nasa-rcmguide.pdf.
Frequently asked questions
RCM is a structured process for determining maintenance policies based on asset functions, functional failures, failure modes, effects, consequences, and the effectiveness of possible tasks.
No. FMEA identifies failure modes and effects; RCM uses this type of analysis within a broader logic to evaluate consequences and select maintenance policies.
No. The outcome may be CBM, scheduled restoration or replacement, failure finding, planned corrective maintenance, redesign, or another technically justified action.
It is a failure of a protective or standby function that may remain unnoticed during normal operation and become evident only when the function is demanded or tested.
Primarily in critical, complex, or high-impact systems where justified maintenance policies can reduce risk, downtime, and lifecycle cost.
Related technical materials
Related services
- Reliability and Availability Engineering: Criticality, Failures, Performance, and Continuity
- Maintenance Engineering: Strategies, Reliability, Plans, and Indicators
- Engineering Asset Management: Registers, Criticality, Lifecycle, and Performance
Related solutions
- Requirements, Evidence, and Acceptance Criteria Management
- Technical Knowledge Management and Lessons Learned
Main content on this topic
- FMEA in Engineering: How to Analyze Failure Modes, Effects, and Causes
- FMECA: How to Analyze Failure Modes, Effects, and Criticality
