Predictive Maintenance Works, but Not for the Reason Most Plants Expect

Predictive Maintenance Works, but Not for the Reason Most Plants Expect

A plant manager once described his predictive maintenance rollout to me as the most expensive way his company had ever learned something it already knew. Two hundred sensors, a dashboard that glowed reassuringly in the control room, and a gearbox that still failed on a Tuesday morning in February. The system had flagged it. Eleven days earlier. Nobody had done anything about the alert, because nobody owned it.

That story is closer to typical than the vendor case studies suggest. The technology behind condition-based maintenance is mature and it genuinely works. The reason programmes disappoint is almost never the physics.

Three Ways to Maintain a Machine

Reactive maintenance means running equipment until it breaks. It is cheap right up until it is catastrophic, and it remains surprisingly common for non-critical assets, which is a defensible choice rather than a failure of nerve.

Preventive maintenance replaces parts on a schedule. Every 2,000 hours, every quarter, whatever the manual says. It is predictable and easy to plan around, and it has one large flaw: most components do not fail on a timetable. You end up throwing away serviceable parts while occasionally still missing the one that failed early.

Predictive maintenance measures the actual condition of the asset and intervenes when the evidence says intervention is needed. Framing this as predictive maintenance vs preventive maintenance sets up a contest that does not really exist in a working plant. Nearly every good programme runs both, with condition monitoring reserved for assets where failure is expensive and calendar-based work kept for the rest.

What Condition Monitoring Actually Watches

There is no single sensor that sees everything, which is why mature programmes layer several techniques.

Infrared thermography finds loose electrical connections, overloaded circuits and failing motor windings, because almost everything that is about to fail electrically gets warm first. Oil analysis reads wear particles, viscosity change and contamination, and it will tell you a bearing is shedding metal long before anything is audible. Ultrasound picks up the very earliest stages of lubrication failure, along with compressed air leaks and electrical arcing. Motor current signature analysis infers mechanical trouble from electrical behaviour without anyone touching the machine.

Each of these catches a different failure mode at a different point on the curve between the first detectable symptom and the actual breakdown. Choosing techniques is really about choosing how much warning you need.

Why Vibration Analysis Is Still the Backbone

For rotating equipment, vibration analysis remains the most informative single measurement available. Imbalance, misalignment, mechanical looseness and bearing degradation each produce characteristic signatures at predictable frequencies, which means a competent analyst can often name the fault, not merely flag that one exists.

Bearing defects are the clearest example. A spalled outer race generates energy at a frequency derived from the geometry of the bearing itself, so the spectrum tells you which element is damaged. Severity thresholds are covered by international standards, and the broader discipline of predictive maintenance has grown up largely around this one technique.

The catch is interpretation. Wireless sensors have made data collection nearly free, and plenty of plants now drown in spectra that nobody is qualified to read. Engineers comparing notes in communities like r/PLC raise this repeatedly: the hardware budget gets approved and the analyst headcount does not.

The Failure Mode Nobody Budgets For

Which brings us back to that gearbox. The most common way these programmes fail is organisational. An alert fires, it lands in a shared inbox, and there is no defined route from that alert to a work order, a spare part and a scheduled window.

The fix is unglamorous. Every monitored asset needs a named owner, a documented threshold, and a pre-agreed action for each severity level. If a high alarm does not automatically create a work order in the maintenance system, the programme is a very expensive monitoring exercise rather than a maintenance strategy.

Baselines matter just as much. A vibration reading in isolation means very little. What you are looking for is deviation from that specific machine's normal behaviour, which means capturing a healthy baseline before anything else, and recapturing it after every significant repair.

Do Not Over-Maintain

One counterintuitive result from reliability studies is that a large share of failures occur shortly after maintenance. Reassembly errors, contamination introduced during work, incorrect torque, the wrong lubricant. Every time you open a machine you are running a small risk of breaking it.

That is a serious argument for condition-based intervention. If the evidence says the asset is healthy, leaving it alone is not laziness. It is a legitimate reliability decision, and one that preventive schedules structurally cannot make.

The Part That Gets Skipped

Programmes also stall on documentation, particularly in plants running imported equipment. Alarm thresholds, lubrication specifications and fault codes all live in the manufacturer's manuals, and if those exist only in the supplier's language, the maintenance team ends up guessing or relying on a rough machine translation of a safety-critical procedure.

This is why professional technical translation is a reliability issue and not an administrative one, and why exporters increasingly treat technical documentation as part of the product rather than paperwork attached to it. A sensor that detects a fault eleven days early buys you nothing if the person receiving the alert cannot read what to do next.