Resources Data and AI

Predictive maintenance without history is guesswork

A model can see the signal before a failure. It cannot see what your crew did the last three times.

A predictive maintenance model can watch a bearing’s vibration reading drift and tell you the bearing is wearing. What it cannot tell you is that this asset drifted the same way three times before, that it was never the bearing, and that the crew settled it in forty minutes by resetting a valve two steps upstream. The model has readings. It has no record of what happened last time. That gap is why so much predictive maintenance in manufacturing is correct in general and wrong about your plant.

The most uncomfortable meeting in machine learning work is the model review. One I still think about went like this. The model’s warning score on a process pump had been climbing for eleven days. The output said bearing wear, high confidence, recommend replacement at the next window. The reliability lead let it sit for a second and then said, that is the discharge valve. It does this every winter when we run the heavier blend. We have chased it three times and replaced two perfectly good bearings doing it.

He was right. He had also never written any of that down, because he had never needed to. He remembered.

Why a predictive maintenance model gives you a generic answer

Predictive maintenance uses condition data from equipment, such as vibration, temperature or current draw, to estimate when an asset is heading toward failure so the work can be planned instead of reacted to. A model learns to match a sensor pattern to a name for what went wrong. The signal comes from your sensors and is usually fine. The label comes from your maintenance system, and that is where the trouble starts.

Work orders record what was replaced. They rarely record what caused the failure. So a closed job reads “replaced bearing, pump 4,” and the model faithfully learns that this sensor pattern means a bearing. Three of those winter events are now sitting in the training set as bearing failures. The model is not wrong about the data it was given. The data is a record of the shop’s paperwork, not a record of what actually failed.

That is the mistake worth watching for. The dangerous output is not an obvious mistake, which somebody catches. It is a plausible, confident recommendation that fits equipment in general and does not fit yours, and it costs you a good bearing and a shutdown window before anyone notices.

There is a second version of the same problem. Your operation already solved some of these failures. The fix was clever, local, and specific: a warm-up sequence, a torque spec nobody changed in the manual, an order in which two valves have to be opened. None of it is in the data. A model trained on your failures but not on your fixes will keep recommending the answer you rejected in 2022.

What a failure record has to contain

A failure event has to hold six things before it is worth anything to a model, or to the new hire reading it in three years. Only two of them are automatic.

  1. The condition data before the event. At a resolution that still shows the early signal, with timestamps lined up across the systems it came from. If the historian’s data compression threw away the two-second spike, the evidence is already gone.
  2. The label, at the part level. What actually failed. Not “pump down.” The bearing, the seal, the coupling, the valve seat, the controller.
  3. The conditions. What the operation was doing. Product or recipe, load, ambient, upstream state, how long since the last start. Most failure modes are conditional, and a model that cannot see the condition cannot see the mode.
  4. The intervention, in order, including what did not work. The two things tried first that failed are more informative than the one that worked. Nobody logs them. They are the most valuable line in the record.
  5. The outcome. Did it hold, and for how long. A fix that failed again in nine days is a different event from one that held for two years, and right now both close the same way.
  6. What the experienced person thought it was. In their own words, attached to the same timestamp, even when it contradicts the work order. This is the one operations skip, and it is the one that carries the judgement.

Two of these come free from your systems. Four require a person to write a sentence at the moment they still remember. That is the whole cost, and it is the reason most operations do not have history. Nobody decided against it. It was never anybody’s job for the ninety seconds it takes.

How to audit your failure history in an hour

Pick one asset that matters. Pull its last three failures. Try to fill in those six fields using only systems, no phone calls.

Time it. Note every point where you had to walk over and ask somebody. Then look at what you produced and ask a plainer question: could a competent engineer who joined last month diagnose the fourth failure from this? If the answer is no, no model will do better, because the model is reading the same file.

Most people who run this find the same three gaps. The condition data exists but the timestamps do not line up across sources. The labels name a part instead of a cause. And the interventions that did not work were never recorded anywhere.

The fixes are small. Add one required free-text field at work order close-out that asks what was tried and what happened, and keep it short enough that people actually fill it. Stop letting the part number stand in for the cause. Line up the timestamps across your systems once and write down what you did. None of that is a project. All of it changes what you can train on next year.

Why more sensors will not fix this

The usual response to a disappointing model is more instrumentation. Sometimes that is right. Usually it is not.

More sensors give you more of what you already have plenty of: a better description of the present. They do not tell you what this pattern meant the last three times or what the crew did about it. In most plants I have walked into, the limit is not how finely the sensors sample. It is that the failures the plant has already lived through were never written down in a form anything can learn from. Adding sensors gives a sharper picture of the present and still no record of the past.

Your equipment is producing failure records a model could learn from right now. Every failure this quarter is one, and each one is either captured with enough around it to be useful or thrown out at close-out. The models will keep getting better on their own. The record will not. Whatever your predictive maintenance program is worth in three years was decided by what somebody wrote down this week, and by whether the person who knows why the pump does that in January is still here to be asked.

Questions people ask

Predictive maintenance uses condition data from equipment, such as vibration, temperature, current draw or pressure, to estimate when an asset is heading toward failure so the work can be planned instead of reacted to. It sits between run-to-failure and fixed-interval preventive maintenance.