Skip to main content

Industrial digital economy

Predictive maintenance fails on a shortage of failures

Machine learningEmerging TechPublished

A predictive maintenance project usually begins with a reassuring number: terabytes of sensor history, years of it, sampled fast. It then stalls on a much smaller number, which is how many times the failure being predicted has actually occurred on that asset. Models learn from examples of the thing. Well-maintained plants are, by design, places where the thing hardly ever happens.

This is the structural awkwardness at the centre of industrial machine learning, and it is not solved by more sensors. A bearing that fails twice in a decade gives a supervised model two labelled examples. A fleet of forty identical bearings gives rather more, but only if the forty are genuinely identical in duty, mounting and environment, which on a real site they are not. Meanwhile the healthy data accumulates at full sampling rate, producing a dataset that is enormous and almost entirely composed of the uninteresting class.

The asset with the best maintenance record is the asset with the least to learn from. Teams discover this after the data warehouse is built.

Where the labels actually come from

The labels for a predictive model are not in the sensor stream. They are in the maintenance management system: work orders, fault codes, the free-text note a fitter typed at two in the morning. That is the ground truth, and its quality determines the ceiling of the project. A work order that says «replaced bearing» with a timestamp from the end of the shift rather than the moment of intervention is a label with hours of uncertainty attached. A note that says «fixed» is not a label at all.

Which means the first piece of work on most predictive maintenance projects is historical reconstruction: reading the work orders, interviewing the people who wrote them, and establishing when each event began rather than when it was recorded. This is slow, it looks nothing like data science, and skipping it produces a model trained on administrative timestamps.

The frameworks that already exist

Condition monitoring was a discipline long before anybody put a model on it, and its framework documents are worth reading precisely because they are not excited. ISO 17359 sets out general guidelines for condition monitoring and diagnostics of machines, including the unfashionable step of deciding which failure modes matter before choosing what to measure. ISO 13374 addresses the data processing and presentation side, separating acquisition, manipulation, state detection, health assessment and prognosis into distinct layers.

That layering is the useful part. Most failed projects collapse several of those layers into one model and then cannot say which layer is wrong when the output is useless. A pipeline with explicit state detection can be debugged. A single end-to-end predictor cannot, and on a shop floor an undebuggable warning is an ignored warning within about three weeks.

What works when labels are scarce

The honest approaches do not pretend to have examples they lack. Anomaly detection learns the healthy envelope and flags departure from it, which needs no failure examples and produces no prediction of what will fail — a trade that is often acceptable. Physics-based models, where the failure mechanism is understood, carry their own structure and need data only to fit parameters. Condition indicators derived from the mechanism, rather than learned, remain the most reliable thing in rotating machinery and have for decades.

Where a supervised model is genuinely the answer, the realistic route is fleet-wide pooling with an explicit account of what makes the units comparable, and acceptance that the model will be retrained as examples accrue rather than delivered finished.

The question to ask at the start

Before any of this: what decision changes if the prediction is right? A warning two weeks before a failure is valuable if the spare has a two-week lead time and the line can be stopped in a planned window. It is worth nothing if the intervention requires an outage that is scheduled annually regardless. The number of industrial machine learning projects that are technically competent and operationally pointless is not small, and the question that would have caught them takes one meeting.

None of which is an argument against the field. It is an argument for spending the first month in the maintenance records rather than in the model, and for saying out loud, early, how many examples of the target event exist. If the answer is three, the project is not a prediction project yet, and saying so is cheaper than discovering it in month nine.