Root Cause Analysis: A Short Method for When There Is No Time for the Long One
Formal root cause analysis needs a room, a facilitator and half a day. Most failures never clear that bar, so they get fixed, closed and repeated. Here is a twenty minute version.

Everyone agrees root cause analysis is worth doing. Far fewer plants actually do it, and the reason is not laziness.
The formal methods need a room, a facilitator, several people for half a day, and a failure big enough to justify all that. Most failures do not clear that bar. So they get fixed, closed, and repeated three months later.
What follows is a shorter pass. It is not a replacement for a proper investigation on a serious event. It is what to do on the ordinary failures, where the realistic alternative is doing nothing.
Five questions, twenty minutes

You can do this standing at the machine with the person who fixed it, while the evidence is still on the bench.
That timing matters more than the method. A conversation next to the failed part, on the day, gets you further than a formal session three weeks later. By then the component has been thrown away and everyone is working from memory.
Name the component, not the machine
The first question sounds trivial and it is where most attempts go wrong.
The pump failed is not a finding. The mechanical seal failed is a finding, because it points at something specific. If your failure records are full of the first kind, you have a record of what stopped, not a record of what broke.
Be equally specific about the mechanism. Not the seal was damaged, but the seal faces overheated. Damage is what you observed. Overheating is what happened.
The question everyone skips
Question four is why nothing caught it, and it is usually the most valuable one in the list.
Most failures give warning. Something ran hot, sounded different, leaked slightly, drew more current. If it failed without warning, that is worth knowing too, because it changes what maintenance strategy makes sense for that asset.
So ask what would have caught this. Was there an inspection that should have found it, and did it happen? Was there a reading that would have shown it, and does anyone look at that reading? Was the warning there and ignored because it did not seem serious?
The answer often points at a fix that is cheaper than anything mechanical. An inspection moved earlier in the round. A limit somebody actually checks. A note in the job plan.
Know when to stop
The classic mistake is going too deep. Ask why enough times and every failure ends at training, culture or budget. Those answers are true, and you cannot do anything with them this week.
Stop when you reach something you can actually change. That is usually two or three steps in.
The bearing failed because it was not lubricated, because the lubrication route skips that machine when the guard is bolted on. You can fix the route or the guard. That is a real action. Continuing to why maintenance is under resourced is a different conversation with different people.
One action, not five
Finish with a single change, a named owner and a date.
The temptation is to list everything you noticed. A list of five actions from a twenty minute conversation is a list where nothing gets done, because nobody owns any of it strongly enough.
Pick the one that most reduces the chance of a repeat. Write the others in the record if they matter, but do not pretend they are commitments.
Write it where someone will find it
Four or five lines on the work order is enough. What broke, how, what let it happen, what you changed.
That short note is worth a great deal six months later. The same asset acts up, and somebody needs to know whether this has happened before. It is also what turns a pile of failure records into something you can study.
If your records currently say replaced pump and nothing else, adding these four lines is the single biggest improvement available to your failure history.
Aim it at repeats
Do not do this on everything. You will not sustain it, and most failures do not warrant it.
Do it on the ones that keep coming back. Pull the assets that failed more than twice this year and work through those.
Repeats are where a short pass pays back most, because you already know the failure will happen again. Twenty minutes spent on something that recurs quarterly is a much better trade than an hour on a one off.
When to escalate
Some failures deserve the full treatment, and it is worth being clear about which.
Anything with a safety consequence. Anything with a large production loss. Anything where the short pass leaves you unsure. And anything that keeps repeating after you have applied a fix.
That last one matters. If the short pass gave you an answer and the failure came back anyway, the answer was wrong. That is the moment to get people in a room properly.


