Resources Manufacturing

How to run root cause analysis from records

Root cause analysis in manufacturing stalls when the only answer is what one person remembers. Here is how to work from records instead.

Most root cause analysis in manufacturing stalls at the same moment: the step where the next answer depends on what somebody remembers. Start from the record instead. Rebuild the hour before the failure out of what your systems already wrote down, the alarms, the parameter edits, the lot numbers, the work orders, who was on the line, and only then bring people in to explain that timeline. Memory is where an investigation should finish. It is a terrible place to start.

I spent a night shift once on a filler that was rejecting seals in bursts. The next morning, twenty minutes into the meeting, a millwright with twenty-two years on that line said it was the second pallet from the new supplier and that we had seen the same thing in the spring. He was right. The meeting was over in four minutes and everybody left happy. I was not happy. We had not solved anything in four minutes. We had gotten lucky, because the only record of that finding was in one man’s memory. The next time it came back he was on vacation, and the same failure ran for most of two weeks.

Why root cause analysis in manufacturing stalls

Root cause analysis is the search for the change or condition that produced a failure, as opposed to the symptom that tripped the alarm. Recall is fast, confident and impossible to check. That is a bad combination in an investigation. Nobody in the room can tell the difference between a veteran who remembers the actual cause and a veteran who remembers a plausible one. A 5 whys session run on recall settles on whoever sounds surest.

The record almost always exists. It is just scattered. The historian has the alarm sequence and the trend. The maintenance system has the work order where someone changed a bearing at 02:40. The quality system has the lot. Purchasing has the receipt. The shift log has three lines of handwriting. Nobody has ever put those five sources on one timeline, so memory wins by default. It is the only version of events that fits inside a one-hour meeting.

Four kinds of records to pull

  • Machine. Alarms and faults in order, cycle times, setpoints and every edit to them, and the state of the equipment upstream and downstream. The fault that trips first is often a symptom of something that drifted twenty minutes upstream.
  • Material. Lot and batch numbers, receipt dates, supplier, moisture or gauge or temper, and where else that lot ran. A cause that follows the material will show up on a different line on a different shift, which is the fastest confirmation you will ever get.
  • People. Who was running it, which shift, first day back, sixth day straight, a new operator on a machine that still takes practice to run well. This is not about blame. It is how you catch the failure that happens on Sundays and never on Tuesdays.
  • Changes. Work orders closed, preventive maintenance done or skipped, recipe and parameter edits, a part swapped for an equivalent, a program updated, a setpoint somebody nudged at 3am. Failures rarely start on their own. Something changed, and the change is usually the shortest path to the cause.

How to rebuild the timeline from records

A spreadsheet is enough for every step below.

  1. Start at the last shift that ran clean. Do not open the timeline at the moment of failure. Open it at the last shift the line ran clean and walk forward.
  2. Put everything on one timeline, in one file. A timestamp column, a source column, a description column. Machine events, material moves, staffing, changes, sorted together.
  3. Mark every change, including the routine ones. A completed PM is a change. A supplier substitution approved three weeks ago is a change. Colour them differently and read the timeline again.
  4. Write down what would have to be true. For each candidate cause, name the evidence that would prove it and the evidence that would kill it. Then go get both. This is the step that separates an investigation from a discussion.
  5. Bring people in to explain the gaps. Once the timeline exists, the operator with eleven years on that line becomes the most valuable person in the building, because now he is interpreting evidence instead of supplying it. Ask him what the record does not show. That answer is usually the real finding.
  6. File the answer where the next investigation will look. Attach it to the reason code, the asset, the part number, the supplier lot. Something that will surface on its own the next time the same conditions line up.

Try step one and two on last month’s biggest downtime event. Twenty minutes, records only, no phone calls. Whatever you cannot reconstruct is not a gap in your investigation skills. It is the list of things your plant does not currently write down, and it is worth more than the investigation itself.

Why the same root cause keeps coming back

Because the finding got written into a report, and the report went into a folder named by date, and nobody searches a folder they do not know exists. Six months later a different crew in a different building spends the same days finding the same answer.

Every shift produces findings like this. A near miss caught early. A cause finally nailed down. A workaround that saved a changeover. Most plants let them disappear at the end of the shift. Captured properly, they are the most valuable thing the operation makes, because each one shortens the next investigation. The test of a captured learning is simple: does it show up on its own the next time, or do you have to remember to go look for it? A finding attached to a reason code finds you. A finding in a PDF does not.

Every investigation that ends because somebody remembered depends on that person still being there. Eventually he retires, takes a vacation, or moves to days. Then the same failure comes back on a night shift, on a line that is down, in front of people who were not there the first time. You can wait for that, or you can spend twenty minutes this week finding out what your plant can still remember without him.

Questions people ask

Root cause analysis in manufacturing is a structured attempt to find the change or condition that produced a failure, rather than the symptom that tripped the alarm. In practice it means rebuilding what happened before the failure from machine, material, staffing and change records, then testing each candidate cause against that evidence.