Bad actor analysis for maintenance and reliability

Find the assets consuming your maintenance resources.

Bad actor analysis uses maintenance history to identify equipment and failure patterns that consume a disproportionate share of labor, downtime, cost, emergency work, or management attention. MaintenanceAI turns that history into a ranked list with supporting evidence and a clear next action.

Start with the free Health Check See sample bad-actor output

Rank by evidence, not by whichever breakdown was loudest last week.

A bad actor can be bad in different ways
Failure frequencyRepeated corrective or breakdown events on the same asset or failure mode.
Downtime impactFailures that stop or constrain production for long periods, even if they happen less often.
Maintenance burdenAssets consuming excessive labor, parts, emergency response, or repeat troubleshooting.
Business consequenceSafety, quality, environmental, throughput, or customer impact that elevates an asset beyond simple work-order count.
What is bad actor analysis?

It is a targeting method for reliability work.

A maintenance team cannot perform a deep reliability investigation on every asset at the same time. Bad actor analysis narrows the field by showing which assets, systems, or failure modes are creating the largest reliability loss or maintenance burden.

Frequency

How often is the asset failing, generating corrective work, or returning with the same symptom after repair?

Impact

How much downtime, production loss, labor, cost, emergency work, or operational disruption does the asset create?

Recurrence

Are multiple work orders actually the same unresolved failure mode repeating under different descriptions, technicians, or codes?

The purpose is selection: determine where defect elimination, root cause analysis, PM changes, condition monitoring, design correction, or operating changes will produce the highest value.
How bad actors should be ranked

Do not rely on failure count alone.

A pump that fails ten times for ten minutes each can be less important than a compressor that fails twice and stops production for two days. A useful ranking looks across several dimensions and keeps the underlying evidence visible.

Corrective work count

Number of failure-related, breakdown, or corrective work orders in the selected period.

Reactive labor

Actual labor consumed responding to unplanned failures, repeat repairs, troubleshooting, and emergency work.

Downtime

Total unavailable or production-impacting time associated with the asset where trustworthy downtime data exists.

Cost

Maintenance cost, parts usage, contractor spend, or other economic burden where the CMMS captures it reliably.

Repeat failure pattern

Evidence that the same component, symptom, cause, or failure mode keeps returning after previous interventions.

Criticality and consequence

Safety, environmental, quality, throughput, redundancy, and business consequence used to prevent a low-frequency but high-risk asset from being ignored.

Data used in bad actor analysis

Your existing CMMS history is usually the starting point.

Primary maintenance evidence

  • Work-order count and work type
  • Asset and equipment identifiers
  • Failure, problem, cause, and remedy codes
  • Technician closeout notes
  • Actual and planned labor
  • Parts usage and repair cost where available
  • Emergency or reactive classifications
  • Created, started, and completed dates
  • Downtime or production-loss fields where available

Context that strengthens the ranking

  • Asset criticality
  • Production line or process area
  • Redundancy and standby configuration
  • Operating duty and service
  • Known failure modes
  • PM and PdM coverage
  • Condition-monitoring findings
  • Operator and technician observations
  • Existing RCA or engineering investigations

See how MaintenanceAI analyzes CMMS history →

Bad actor analysis workflow

From work-order history to a reliability priority list.

1. Define the period and scope

Select the site, area, line, asset population, and analysis window so comparisons are meaningful.

2. Normalize the records

Standardize asset identifiers, work types, failure codes, dates, and other fields before ranking the population.

3. Build multiple rankings

Rank by frequency, downtime, labor, cost, emergency work, and recurrence rather than forcing every decision into one metric.

4. Cluster repeat failures

Review descriptions, codes, and technician notes to determine whether apparently separate events represent one recurring failure mechanism.

5. Apply consequence and context

Use criticality, production impact, safety, quality, environmental risk, and redundancy to keep the ranking aligned with the business.

6. Assign the next reliability action

Move each high-priority actor toward RCA, PM optimization, condition monitoring, design correction, operating change, or another defined action.

Bad actors and repeat failures

A bad actor list is only useful if it changes the maintenance response.

Weak response

  • Rank the top ten assets once a month
  • Discuss the list in a meeting
  • Repair the next failure the same way as the last one
  • Close the work order
  • Leave the underlying mechanism unresolved

Strong response

  • Validate the recurring pattern
  • Separate symptom from failure mode and cause
  • Assign an owner and investigation path
  • Change the maintenance, design, operating, or work-management system as needed
  • Track whether the intervention actually reduces recurrence
Selection is only the first step. Bad actor analysis tells you where to investigate. Root cause analysis, defect elimination, strategy changes, and disciplined follow-up determine whether the reliability loss actually goes away.
Data quality matters

A bad actor ranking is only as trustworthy as the records behind it.

Missing asset IDs, vague closeout notes, inconsistent failure codes, unreliable downtime, and incomplete labor capture can distort rankings. MaintenanceAI keeps those limitations visible instead of hiding them behind a single score.

Asset identity

If the same machine appears under several names or locations, its failure burden can be split across records and disappear from the top of the list.

Failure definition

If every problem is coded as "mechanical" or "other," the asset may rank correctly while the recurring failure mechanism remains invisible.

Impact capture

Downtime, labor, and cost should be used only to the extent the site records them consistently enough to support the decision.

What happens after the ranking?

Every top bad actor should have a disposition.

Root cause investigation

Use when the recurring failure mechanism is consequential, unresolved, and worth deeper technical investigation.

PM optimization

Use when current preventive work is missing the failure mode, poorly written, duplicated, or not producing useful findings.

See the PM optimization method →

Condition monitoring

Use when deterioration can be detected more effectively through condition-based or predictive methods than by fixed-interval intrusive work.

Design or installation correction

Use when chronic failures point to alignment, contamination, piping strain, foundation, component selection, accessibility, or another engineered defect.

Operating correction

Use when the equipment is being run outside the intended operating envelope, duty, startup method, or process condition.

Work-management correction

Use when the apparent reliability problem is being amplified by poor planning, missing parts, bad job plans, delayed corrective work, or weak closeout discipline.

MaintenanceAI Bad Actor / Repeat Failure Analysis

Make the priority list defensible.

Included in the Maintenance Intelligence Assessment

The $1,500 founding-customer assessment includes a focused bad-actor and repeat-failure analysis for one facility and a representative production area or approximately 25 priority assets.

Typical turnaround: about 10 business days after usable data is received.

Start with the free Health Check

What you receive

  • Ranked bad-actor list
  • Supporting frequency, labor, downtime, and other available evidence
  • Repeat-failure patterns
  • Data-quality limitations and confidence
  • Recommended next reliability action
  • Owner-ready action register
  • Connection to PM, backlog, planning, and maintenance-system findings

See the full assessment scope →

Common bad actor questions

What reliability teams usually need to decide first.

How often should we update the bad actor list?

The review cadence should match how quickly the plant’s failure history changes. Many teams review monthly or as part of a regular reliability meeting, but the important point is consistent ownership and follow-through.

Should we use a single weighted bad-actor score?

A weighted score can be useful when the weighting is transparent and agreed. MaintenanceAI also keeps the underlying rankings visible because a single score can hide whether an asset is high because of frequency, downtime, cost, or consequence.

Is the top bad actor automatically the first RCA?

Not always. The team should consider business consequence, whether the pattern is validated, whether the cause is already known, and whether a simpler corrective action can remove the defect before launching a deeper investigation.

Can we do bad actor analysis with poor failure codes?

Yes, but confidence will vary. Work-order descriptions, labor, asset history, downtime, and technician notes can still reveal patterns. Poor coding should also become a specific data-quality improvement action.

MaintenanceAI

Stop fixing the same assets without changing the system behind the failure.

Use CMMS history to identify which assets are consuming your maintenance capacity, understand why they keep returning, and assign the reliability actions that deserve attention first.

Start free Health Check See sample bad-actor output