Resources
/
Maintenance and Reliability Metrics
Maintenance and Reliability Metrics

Mean time to repair (MTTR): how fast your plant gets back on its feet

A centrifugal pump trips at six in the morning. By the time someone notices it is down, calls out a technician, diagnoses the fault, hunts down the right seal, finishes the repair and hands the pump back to operations, it is nearly midday. Six hou...

4 min read
On this page

A centrifugal pump trips at six in the morning. By the time someone notices it is down, calls out a technician, diagnoses the fault, hunts down the right seal, finishes the repair and hands the pump back to operations, it is nearly midday. Six hours for a job that, with the wrench actually turning, takes two.

Those four extra hours are not just delay. They are a verdict on how well your whole repair system works, and that is exactly what mean time to repair measures.

What it actually measures

MTTR is the average time it takes to get a failed asset back to full working order. It is a measure of maintainability, sometimes written as mean time to repair or replace, since for some items the quickest route back is a swap rather than a fix. Crucially, it covers far more than the spanner work: the time to notice the fault and mobilise, to diagnose it, to find the spare, to do and test the repair, and to hand the machine back. The whole restoration cycle, from "it stopped" to "it is running again".

How to work it out

MTTR = Total repair time (hours) / Number of repairs

An asset fails 10 times, with individual repairs of 2, 6, 10, 6, 5, 10, 1, 2, 5 and 3 hours.

MTTR = (2 + 6 + 10 + 6 + 5 + 10 + 1 + 2 + 5 + 3) / 10 = 50 / 10 = 5 hours

The team then kits parts ahead of time and sharpens its diagnosis checklist, and the next quarter the same number of failures is cleared in an average of 3 hours each. The asset is no more reliable than before, but every failure now costs far less production.

Why most of MTTR is waiting

Here is the insight that changes how you improve it: in most plants, the wrench is turning for only a small fraction of the time an asset is being "repaired". The rest is waiting. Waiting for someone to notice and mobilise, waiting for a permit or a lockout, waiting for the right part to be found and carried to the job, waiting for a second trade to arrive. Strictly, the hands-on portion is active repair time; everything around it is logistic and administrative delay. A high MTTR almost never means your people are slow with a spanner. It means the system around them is making them wait.

How to cut it

Once you see MTTR as mostly waiting, the levers are obvious. The biggest is preparation: kitting and pre-staging the parts, tools and instructions for a job before it starts, so the technician walks to a ready job rather than a treasure hunt. The next is design. Equipment built for maintainability, with standard parts, interchangeable and modular components you can swap rather than rebuild in place, and honest access so you are not removing three healthy things to reach the broken one, can collapse a repair time that the crew could never fix by working harder. And then skills: a technician who can diagnose quickly, with standard tools and clear work instructions to hand, spends minutes where a guesser spends hours.

Where it can mislead you

  • Beware the figure that quietly drops the waiting. If your records capture only the active repair and ignore the permits, parts and travel around it, your "MTTR" is really the smaller, flattering core of the true lost time. Be clear about which you are measuring.
  • The spread matters as much as the mean. Repair times are rarely tidy, and a couple of monster events can drag the average somewhere that describes no actual repair.
  • Calculate it by asset or asset class, not just plant-wide. A fleet average smooths over the one chronic bad actor that most needs your attention.
  • Watch it next to mean time between failures. A falling MTTR alongside a falling MTBF is a warning, not a triumph: you are getting better at fixing failures that are happening more and more often.

Mean time between failures tells you how often an asset fails; MTTR tells you how quickly you get back on your feet when it does, and together they make up availability, the share of time the asset is actually there for you. Drive MTTR down and you claw back production without touching the failure rate at all, simply by clearing the path so the hours your people give you reach the job instead of the queue.

Found this useful? Share it with your team.
Share on LinkedIn

Ready to elevate your skills or empower your team?