XRVE Reliability knowledge base | Performance management

The Maintenance KPIs That Actually Matter

The best maintenance dashboard is not the one with the most metrics. It is the one that shows whether the maintenance system is gaining control, whether failure behavior is improving, and where leadership needs to remove a constraint or change the strategy.

Direct answer

Which maintenance KPIs matter most?

A useful maintenance scorecard needs a small set of metrics that cover five questions: Are we controlling incoming reactive work? Are we preparing and executing the weekly schedule? Are PM and PdM strategies actually controlling failure? Is equipment reliability improving? Can the underlying CMMS data be trusted? Cost belongs in the picture too, but cost without workload, consequence, and reliability context is easy to misread.

No single KPI proves maintenance excellence. High PM compliance can coexist with breakdowns. High wrench time can coexist with constant firefighting. Low maintenance cost can be achieved by deferring work until the risk is worse. Metrics should be read as a system.

The XRVE KPI stack

Measure control before celebrating outcomes.

1. Demand controlReactive work, emergency work, repeat failures, and incoming corrective demand.
2. Work managementReady backlog, schedule load, schedule compliance, break-in work, and planning accuracy.
3. PM effectivenessPM compliance plus findings, corrective conversion, failure-mode coverage, and recurrence after PM.
4. Reliability outcomesAsset-level failure frequency, downtime, MTBF/MTBD, repair duration, and availability where definitions are trustworthy.
5. Data confidenceAsset linkage, labor capture, closeout quality, meaningful dates, failure coding, and status discipline.
Joshua's field note

If a KPI cannot trigger a management question or a decision, I do not want it taking prime space on the dashboard. Every metric should tell us either where control is improving, where the system is breaking down, or what assumption needs investigation.

Core maintenance scorecard

A practical set of metrics for most industrial maintenance teams.

KPIWhat it tells youWhat to pair it withCommon trap
Reactive workHow much maintenance labor is being pulled into schedule-interrupting or unplanned response.Emergency work, break-in causes, repeat failures.Calling every corrective job reactive even when it was identified and planned in advance.
Schedule complianceWhether the locked weekly schedule is being completed as committed.Schedule load, break-in work, ready backlog.Improving compliance by underloading the schedule.
Ready backlogWhether enough executable work exists to build a realistic schedule.Total/planned backlog, age, constraints, labor capacity.Treating a CMMS status code as proof of readiness.
PM complianceWhether preventive work is executed inside the defined timing rule.PM effectiveness, repeat failures, finding yield.Assuming on-time completion proves the maintenance strategy works.
PM/PdM effectivenessWhether proactive tasks identify defects and create useful corrective action before failure.Failure-mode coverage, corrective conversion, reliability trend.Forcing effectiveness into one percentage without validating what the numerator means.
Repeat failure burdenHow much labor, downtime, cost, and attention recurrence is consuming.Asset criticality and bad-actor ranking.Counting similar symptoms as identical root causes.
MTBF or MTBDTrend in time or operating exposure between defined failures or downs for a specific population.Failure definition, operating context, criticality.Mixing unlike assets, changing failure definitions, or using fleet averages that hide bad actors.
MTTR / mean downtimeHow long restoration or downtime events take under a defined rule.Planning quality, parts, access, troubleshooting, maintainability.Assuming a long repair automatically means technician performance is the issue.
Maintenance costResource consumption and financial burden.Asset value, production, risk, workload, reliability trend.Treating low spend as success without checking deferred work or reliability loss.
CMMS data qualityWhether the metrics above are trustworthy enough to support decisions.Field completeness, consistency, specificity, validity.Building sophisticated dashboards on weak asset and work-order history.
Leading and lagging indicators

Pair what happened with what the organization can still influence.

Lagging: downtimeLeading partner: ready backlog and high-criticality defect age.
Lagging: repeat failuresLeading partner: RCA action closure and PM strategy changes.
Lagging: emergency workLeading partner: defect identification, planning throughput, and proactive finding conversion.
Lagging: maintenance costLeading partner: backlog risk, planned work, rework, and asset health.
Lagging: MTBFLeading partner: condition findings, critical PM effectiveness, precision maintenance controls.
Lagging: schedule missesLeading partner: job readiness, materials, access, estimate quality, and break-in work.
Three work-management metrics that belong together

Schedule compliance alone is easy to game.

Schedule load

How much of realistic weekly labor capacity was intentionally committed to the schedule?

Schedule compliance

How much of the committed work was completed under the site's defined rule?

Break-in work

How much work entered after the schedule was locked, why did it enter, and what committed work did it displace?

A site can report 100% schedule compliance by scheduling very little. A site can also schedule aggressively and miss the week because high reactive demand keeps breaking in. Read the three signals together, then use ready backlog to test whether planning is giving the scheduler enough options.

Reliability metrics

Use MTBF, MTTR, and availability at the level where the definition is defensible.

Reliability metrics are strongest when the equipment population, failure definition, operating exposure, and data source remain stable. They become weak when a plant mixes unlike assets, records only some failures, or interprets every work order as a functional failure.

MTBF / MTBD

Use for trends on comparable assets or a defined system. Know what counts as a failure or down event and whether operating time or calendar time is appropriate.

Repair or downtime duration

Use to identify maintainability, access, parts, troubleshooting, coordination, or planning constraints. Do not automatically convert the number into a labor-performance judgment.

Availability

Useful when uptime and downtime states are consistently defined and maintenance has meaningful influence over the losses being counted.

What not to manage as a trophy metric

Some metrics are more useful as diagnostics than as permanent targets.

Wrench time

Useful for exposing waiting on parts, permits, access, travel, information, or coordination. High wrench time does not prove strong reliability if technicians are constantly repairing failures.

Work-order count

Useful for workflow and record-volume analysis, but weak for labor capacity because a work order can represent ten minutes or several hundred hours.

PM count completed

Useful for throughput but dangerous without task quality, failure-mode relevance, findings, and follow-up.

Do not make a metric the objective. The objective is reliable, safe, cost-effective operation. Metrics are evidence about whether the maintenance system is moving toward or away from that outcome.
A dashboard for different levels

The maintenance manager and the planner should not need the same dashboard.

Plant / business level

High-consequence failures, availability where valid, maintenance cost context, critical asset health, major backlog risk, and reliability trend.

Maintenance management

Reactive work, emergency work, planned work, schedule compliance, ready backlog, PM effectiveness, repeat failures, rework, and data confidence.

Planner / supervisor level

Ready work by craft, job readiness, aged planned work, material constraints, weekly load, break-in causes, estimate accuracy, and technician feedback.

A 30-minute KPI audit

Remove dashboard noise before adding another metric.

1. List every reported maintenance KPIInclude dashboards, daily boards, weekly reports, corporate scorecards, and local spreadsheets.
2. Write the management questionFor every metric, state the decision or question it is supposed to support.
3. Write the exact definitionDocument numerator, denominator, exclusions, timing rule, population, and source system.
4. Identify the companion metricAsk what second signal is needed to prevent the first KPI from being misread or gamed.
5. Check data confidenceRate whether the underlying fields are complete, consistent, and specific enough to trust the trend.
6. Remove orphan metricsIf nobody changes a decision when the KPI moves, demote or delete it from prime dashboard space.
XRVE operating principle

A KPI should create a better question, not end the conversation. Green numbers deserve investigation when failure behavior, backlog risk, or frontline experience tells a different story.

Technical references

References and further reading

About the author

Joshua Rivera is a maintenance and reliability leader and founder of XRVE Reliability. His experience includes maintenance leadership, planning and scheduling, KPI development, CMMS ownership, PM optimization, failure analysis, troubleshooting, and reliability improvement.

Read Joshua Rivera's background →

XRVE Reliability

Measure the maintenance system, not just the volume of work.

Use a compact scorecard that shows demand, work-management control, PM effectiveness, reliability outcomes, cost context, and data confidence.

Start free Maintenance Health Check Use the free maintenance tools