Guide / Industrial AI
Predictive maintenance: Build a pilot your maintenance team can use
A machine shows an unusual pattern. The replacement part takes five days to arrive, and the next planned shutdown is tomorrow. The useful question is whether the warning supports a better maintenance decision. An alert without enough lead time or a responsible person may add work without changing the outcome.
Predictive maintenance uses condition and operating data to estimate future maintenance needs. For a manufacturing business, a promising starting point is an asset with costly failures, observable deterioration and a practical response to a warning. Those conditions matter more than the sophistication of the model.
Choose the maintenance strategy before the algorithm
Reactive maintenance repairs an asset after failure. Preventive maintenance follows an interval based on time or usage. Condition-based maintenance responds to an observed state. Predictive maintenance adds an estimate of future development. A temperature threshold alone does not establish remaining useful life.
Fraunhofer IESE discusses predictive maintenance and its implementation challenges. Our recommendation for a first project is to work backwards from the maintenance action: identify the decision, the required warning window and the evidence that could support it.
| Asset situation | First approach to consider | Question to answer |
|---|---|---|
| Inexpensive, easily replaced component | Reactive maintenance | Is a failure genuinely low consequence? |
| Well-understood wear pattern | Scheduled maintenance | How often are useful parts replaced early? |
| Observable change in condition | Condition monitoring | Which observations justify an inspection? |
| Repeated failures with a detectable precursor | Prediction pilot | Can the team act before the failure? |
A plant can use all four approaches. A critical spindle and an accessible auxiliary drive need not share the same strategy.
Bound the pilot to one failure mechanism
Consider bearing wear in a family of pumps. Define whether a warning should trigger an additional measurement, an inspection or preparation for a planned replacement. The eventual action remains subject to engineering evidence and the relevant operating approvals.
Vibration, temperature and load may be useful for this example. Other mechanisms may call for current measurements, oil analysis or thermal imaging. Select measurements that relate to the failure mechanism. A large collection of unrelated signals is not a substitute for that reasoning.
Record normal changes as well: startup, partial load, cleaning and product changes. Otherwise the system may treat an acceptable transition as a defect. Maintenance records should identify what was found and replaced. “Machine repaired” is usually too vague to establish a useful training label.
How measurements become a maintenance recommendation
Start with sensor readings and the corresponding operating state. Data preparation checks missing measurements and creates useful features, such as a change in vibration under comparable load. A rule or model then assesses the change. The resulting alert should identify the asset, finding, uncertainty and next inspection step.
The maintenance team investigates and records what it actually found. That feedback supports future evaluation. Automatically treating every alert as confirmed damage would feed errors back into the dataset.
Potential applications include pumps in energy facilities, machine-tool spindles and conveyor drives in logistics. These are possible applications, not Welf project references. Each needs its own failure mechanism, relevant measurements and practical maintenance response. Copying a model alone does not establish that it will transfer between assets.
Inspect the evidence before training
Check that timestamps, asset identifiers, operating states and maintenance events can be joined consistently. Look for sensor replacements, missing intervals and changes to the equipment. Confirm that useful measurements exist before the event, rather than only during the breakdown.
Where confirmed failures are scarce, an anomaly detector can flag departures from known behaviour. Treat those flags as inspection prompts. They do not automatically support a claim about when the machine will fail.
Keep a simple rule-based baseline in the comparison. If a transparent threshold gives the maintenance team enough warning, a more complicated model must justify its additional operating cost.
Test with later data and different assets
Use an independent later period for evaluation. Randomly splitting fragments of the same failure across training and test sets can make performance look stronger than it is. If the system is intended for additional machines, test on assets withheld from development too.
Then run alongside the existing maintenance process. The system produces recommendations, while the maintenance team records what it found and why it acted. This reveals whether an apparently useful statistical signal produces a useful operational decision.
| Measure | What it tells the team |
|---|---|
| Relevant events detected | Which confirmed problems received a timely warning? |
| False alarms per asset per week | How much extra investigation is required? |
| Usable lead time | Was there time to obtain parts and organise work? |
| Missed events | Which failures produced no useful warning? |
| Time spent handling alerts | What does the response process cost? |
Set acceptable alert workload before reviewing the results. A model can detect many events and still be impractical if engineers must investigate too many false alarms.
Calculate value without inventing avoided failures
Include sensors, integration, data preparation, model maintenance and alert handling. Count benefits only where a better maintenance decision can be supported by evidence. An alert is not automatically an avoided breakdown.
For illustration, 20 monthly alerts taking 30 minutes each require ten hours of investigation. If only two lead to useful action, their value must cover the wider workload as well. Whether either action prevented a failure requires a separate engineering assessment. These numbers are hypothetical, not Welf project results.
Lack of room to act is a valid reason to stop. If parts cannot be sourced and the maintenance window cannot change, earlier information may have limited value. Improve the response process before investing further in prediction.
Make the result operable by your team
Agree who investigates sensor faults, reviews rising alert rates and approves new model versions. Equipment modifications and changes in operating conditions may require renewed evaluation. Keep data formats, test cases and alert history accessible so the system can be reassessed later.
A pilot should leave the company with an operating procedure and a reproducible evaluation, not just a dashboard. The principles in our industrial AI model-switch protocol help frame repeatable comparisons. Infrastructure choices are covered in private AI, European cloud or API.
Welf develops and integrates industrial AI around operational decisions. Bring one asset family, a relevant failure mechanism and a description of the available data to the first discussion. Discuss a maintenance application.