Resources
/
Maintenance and Reliability Metrics
Maintenance and Reliability Metrics

Systems covered by criticality analysis: are you maintaining what matters, or just what shouts loudest?

The planner is building next year's budget, and she needs to know where to aim it: which machines deserve the most spares, the closest watch, the tightest inspection routine. The asset register lists more than eighteen hundred systems, and not one...

5 min read
On this page

The planner is building next year's budget, and she needs to know where to aim it: which machines deserve the most spares, the closest watch, the tightest inspection routine. The asset register lists more than eighteen hundred systems, and not one column tells her which of them actually matter.

So, like most plants without a better method, attention drifts to wherever the noise is loudest. The machine that failed last week gets the spares. The manager who shouts gets the inspection. Meanwhile a genuinely dangerous system, quiet for now, gets the same casual treatment as a ventilation fan. Systems covered by criticality analysis is the metric that replaces the noise with a ranking.

What it actually measures

It is the share of your systems for which a formal criticality analysis has been completed and written down. A system here is a set of connected assets with a defined purpose and clear boundaries, not a single component, and "formal" is the operative word: a ranking that lives in one engineer's head does not count.

How to work it out

Systems covered (%) = Systems with a completed criticality analysis / Total systems × 100

A plant with 1,811 systems has completed the analysis for 337 of them.

Systems covered = 337 / 1,811 × 100 = 18.6%

Fewer than one system in five has been ranked. That is not a failure; it is a mandate. The team now knows exactly what to do next: work through the rest, starting with anything suspected to be critical, until risk, not habit, decides where the effort goes.

How the ranking is actually done

Underneath the percentage sits a real piece of analysis, and it always combines the same two ideas: how badly a failure would hurt, and how likely it is. The consequence side is scored across several themes, with safety and the environment carried at the top of the scale and production, quality and cost beneath them; the likelihood side runs from the everyday to the once-in-a-decade. Multiply or combine the two and each system earns a criticality number that then lives against it in the maintenance system and quietly governs everything: how often it is inspected, how many spares it earns, where it sits in the work queue, even whether it makes the cut for a capital project.

The methods range from a simple criticality table through to a full failure modes and effects analysis, and one common scoring approach is worth knowing because of a subtlety it adds. It multiplies three things, not two: how severe the failure is, how often it happens, and how likely you are to catch it coming before it bites. That third factor, detectability, is the one people forget, and it is precisely what condition monitoring improves. The whole exercise is the backbone of risk-based maintenance, the principle that the most resources should flow to the assets carrying the most risk, and the least to the ones that can fail without anyone much caring.

Why it is worth the effort

The payoff is larger than "better prioritisation". Done properly, ranking by risk does not just move effort around; it removes effort that never needed spending. Where this analysis has been applied rigorously, as the foundation of a reliability-centred maintenance programme, it has cut routine maintenance workloads by something like 40 to 70 per cent in documented cases, because so much scheduled work turns out to be aimed at things that do not fail in ways the work would catch. That is the quiet prize hiding behind the coverage figure: not only safer and better-aimed maintenance, but markedly less of it.

Where it can mislead you

  • Pitch the analysis at the right level, or it tells you nothing. Assess a whole aircraft as one thing and you conclude "flying is risky"; drop to the system level and the real picture appears, the structure that almost never fails set apart from the comfort system that often does. Go too low, down to individual components, and you drown; too high, at the whole facility, and you learn nothing. The system level is where it works.
  • Do not lean only on failure frequency. A pure tally of what breaks most often quietly ignores the catastrophe that has never happened yet, the high-consequence, low-frequency event that a good analysis is there to catch. Rank across the whole spectrum of risk, not just the noisy end.
  • Mind the subjective scores. Because the scoring leans on human judgement, especially the detectability factor, a number can be nudged up or down to push an asset across a threshold, so the ranking needs consensus among the people accountable for it, not one person's quiet opinion.
  • Coverage is not quality, and it is not permanent. An analysis done at commissioning and never revisited may no longer match how the plant runs, so assess new systems before they are commissioned and review the rest on a regular cycle, because a system that looked harmless last year may have quietly become essential.

Maintenance budgets are always smaller than the wish list, so the only real question is where the money does the most good. This metric is how you prove your answer is built on risk rather than on whoever happened to shout loudest this week, and the reward for getting it right is a programme that is at once safer, sharper, and a good deal lighter than the one it replaces.

Found this useful? Share it with your team.
Share on LinkedIn

Ready to elevate your skills or empower your team?