Failure-mode-based maintenance

Build PM tasks around how equipment fails, not around how the old PM was written.

Failure-mode-based preventive maintenance asks what can fail, what consequence matters, whether deterioration can be detected, and which maintenance strategy can actually control the risk.

Reviewed and authored by Joshua Rivera, XRVE Reliability founder and maintenance & reliability leader.

Start with failure behavior

The maintenance task should exist because it controls a credible failure mode.

Legacy PM programs often accumulate tasks because equipment was new, a failure happened once, a vendor recommended a generic interval, or someone wanted to “be safe.” Over time the program becomes a mixture of valuable work, vague inspections, duplicates, unsupported frequencies, and tasks disconnected from current asset risk.

A failure-mode-based review reverses the logic. First define the function and the way it can fail. Then choose the task, inspection method, condition threshold, interval, or alternative strategy capable of managing that failure.

Strategy selection

PM is not the only maintenance strategy.

Time or usage based PM

Use when deterioration is meaningfully age- or usage-related and a replacement, restoration, or service interval is technically defensible.

Condition-based / PdM

Use when deterioration can be detected with enough warning to plan corrective action. Examples can include vibration, oil analysis, thermography, ultrasonic inspection, or measured wear.

Failure-finding

Use for hidden protective functions where the failure may not be evident during normal operation and periodic functional testing is needed.

Redesign or engineered control

Use when failure consequence is unacceptable and no feasible maintenance task controls the mechanism well enough.

Run-to-failure

Use deliberately when consequence is low, failure is evident, restoration is straightforward, and proactive maintenance does not create enough value.

Operating / care task

Some failure risk is better controlled through cleaning, lubrication, operating discipline, basic care, setup, or process control rather than a maintenance PM.

Write executable tasks

A technically correct strategy can still fail if the task is vague.

“Inspect motor,” “check pump,” and “service conveyor” do not tell technicians what condition to evaluate or what action to take when the condition is unacceptable.

Task elementQuestion it should answer
Component / pointExactly what is being inspected, measured, tested, serviced, or restored?
MethodHow should the technician perform the task?
Acceptance criteriaWhat does acceptable condition look like, and what authority supports the limit?
ResponseWhat happens when the condition is outside the acceptable range?
Frequency basisWhy is the task performed at this interval?
Safety / accessWhat permits, isolation, guarding, sanitation, or access requirements apply?
Example reasoning

Do not map every failure to “inspect more often.”

Bearing deterioration is detectable

If a critical bearing develops a detectable condition with a usable P-F interval, condition monitoring may provide better warning than frequent intrusive replacement.

Coupling fails from recurring misalignment

Replacing the coupling more often does not control the failure mechanism. The strategy should address alignment, base condition, thermal growth, installation practice, or other verified contributors.

Hidden safety function

If a protective device can fail without becoming obvious during normal operation, periodic functional testing may be more appropriate than visual inspection.

Low-consequence lamp failure

If failure is evident, spares are available, and consequence is negligible, planned replacement may create more work than deliberate run-to-failure.

Govern the PM library

Every task should have a disposition and technical basis.

During PM optimization, label tasks as retain, revise, combine, remove, convert to condition-based, convert to failure-finding, escalate for engineering review, or replace with another maintenance strategy. Record the reason so the program does not drift back into inherited work with no technical basis.

Never invent OEM values or engineering limits to make a task appear complete. If a torque value, lubricant specification, temperature limit, vibration alarm, clearance, or test criterion requires authority, source it or flag it for review.

Related maintenance intelligence

Use failure modes to improve PM content, frequency, and effectiveness.

PM Optimization

Review the full PM library using task quality, history, evidence, and technical authority.

Preventive maintenance optimization →

PM Frequency Optimization

Determine whether task intervals are supported by failure behavior and detection evidence.

PM frequency optimization →

PM Compliance vs Effectiveness

Understand why completing PMs on time does not prove the maintenance strategy is working.

PM compliance vs effectiveness →

XRVE Reliability

Turn maintenance history into a defensible next action.

Use the method yourself, or apply the same reasoning to your facility's CMMS history through XRVE Reliability.

Free maintenance maturity assessment See sample analysis