← All Articles
Maintenance

Predictive Maintenance Pays Only Where the Failure Mode Is Actually Predictable

📅 · 4 min read · Meta Smart Factory Team

Every predictive maintenance proposal has the same shape: sensors on the critical machines, a model that learns the signature of a failure, an alert weeks ahead. It almost never contains the question that decides whether it works — whether the failures actually costing you money give warning at all. That can be answered from your own records, before anyone quotes a sensor.

Reactive, preventive, condition-based and predictive maintenance compared

Before cost enters the argument, consequence does. Anything whose failure has a safety or environmental consequence, and anything under statutory examination — pressure systems, lifting equipment, guards and interlocks, machinery brakes, the proof-test interval on a safety instrumented function — sits outside this analysis, done to schedule whatever the economics say.

What is left is the ordinary population. Reactive means running to failure, legitimate for assets whose failure is cheap, whose spare sits on a shelf, and whose loss does not stop production. The mistake is not having reactive assets, but having ones nobody chose.

Preventive on a calendar or a counter replaces components regardless of condition, and needs no instrumentation beyond a CMMS with strategy-based plans. Two costs go uncounted: the remaining life thrown away on components that were fine, and the fault every intrusion risks introducing — a seal nicked on reassembly, contamination let into a bearing. Increasing preventive frequency on an asset that fails randomly makes availability worse; the airline reliability studies behind reliability-centred maintenance found most failure modes show no age-related wear-out at all.

Condition-based maintenance measures, then acts on a threshold: monthly route-based vibration readings, a quarterly thermographic survey, an oil sample per sump, with no model anywhere in it. It costs a trained person's time and an instrument, delivers much of what people attribute to predictive maintenance, and for many plants is the destination, not a stepping stone.

Predictive adds continuous measurement and a projection forward — a trend extrapolated to a threshold, or an estimate of remaining useful life — plus instrumentation and someone to own the models. That is a gain in coverage and cadence more than in insight, earning its cost when an asset needs warning between route intervals, or when there are too many assets to walk.

What predictive maintenance can detect, and what it cannot

A failure mode is predictable when it degrades progressively, emits something measurable while doing so, over an interval long enough to act in. Reliability-centred maintenance calls that the P-F interval: from the point a failure becomes detectable to the point the asset stops doing its job.

One class of failure defeats the question before it is asked. A protective device — a relief valve, a trip, a standby pump, an interlock — has a hidden function: its failure emits nothing, because nothing depends on it until the event it guards against arrives. You cannot condition-monitor a function that is not being exercised, so you proof-test it, on an interval derived from the protected failure's rate and the risk you will carry, and you record the result.

The modes that behave well are mostly mechanical and mostly rotating. A rolling element bearing spall propagates at frequencies set by the bearing geometry and shaft speed, typically over weeks to months on a well-lubricated bearing at moderate speed and load, and over days or less once lubrication is lost or under high speed and overload. Set the monitoring interval from the shortest credible P-F interval, not the typical one. Gear wear and pitting progress similarly; misalignment and imbalance rarely fail as such but destroy what is around them, visible far in advance. So are belt wear, pump cavitation, seal leakage, heat exchanger fouling, and tooling wear where the process signal moves with the tool.

Most of the rest behave badly. Logic boards, I/O cards and most discrete sensors fail with no gradual precursor and a hazard rate close to random. Operator damage, collisions, setup errors and contamination in the raw material are events rather than processes, and control faults leave nothing in a vibration spectrum. Some mechanical failures have a P-F interval of seconds: a fastener fracture, a sudden overload.

The electrical exceptions are worth naming, because they are where an electrical programme pays. Electrolytic capacitors in DC buses and power supplies dry out progressively with temperature and show it in capacitance, equivalent series resistance or bus ripple, which is why drive manufacturers publish replacement intervals. Most modern drives already trend heatsink temperature and fan condition. Motor insulation degrades in a way insulation resistance, polarisation index and partial discharge trending can follow. Continuous-flex cable in drag chains and robot dress packs is rated in flex cycles, which makes it one of the most schedulable items in a cell.

A described symptom is strong evidence that a usable P-F interval exists: if a technician can say what a failure sounds like a week out, an instrument will see it earlier. The absence of one is much weaker evidence, because the human sensing range is narrow. Nobody hears 40 kHz, nobody feels the impacts envelope analysis extracts before the overall level moves, nobody notices a two per cent torque drift. Where no symptom is described, ask whether the mode puts energy outside human range — ultrasonic, high-frequency vibration, slow process drift — before concluding there is nothing to find.

How to tell if predictive maintenance will work before you buy sensors

This is a records exercise, not a procurement exercise. Take twelve months of unplanned stops on the candidate assets and classify each into progressive mechanical degradation, lubrication, sudden mechanical, electrical, control and software, operator or setup error, maintenance-induced, material, and external supply, with the downtime minutes and cost attached.

Two of those are the ones people leave out, and the two most likely to change the conclusion. Maintenance-induced means a failure traceable to the previous intervention on that asset; the CMMS already records what work preceded the stop, so it can be coded retrospectively, and without the category every fault a preventive intrusion introduced lands in sudden mechanical and disappears. Without a lubrication category, a bearing killed by the wrong grease or a missed route reads as progressive mechanical degradation and looks like a case for buying sensors.

Then read the split, because it decides the project. If the money concentrates in progressive mechanical degradation, predictive maintenance has something to work on — but before quoting an accelerometer, audit the two things that remove those failures rather than predicting them. Lubrication first: correct product, correct quantity, correct interval, contamination control, ultrasonic-guided regreasing instead of a grease gun on a calendar. Installation precision second: laser alignment, soft foot, pipe strain, balancing. Both are cheaper than monitoring, and both compete directly with the sensor budget.

If the money concentrates in operator damage and setup error, the intervention is fixturing, training and mistake-proofing, and no accelerometer touches it. If it concentrates in electronics, the answer is mostly spares on the shelf, redundancy on critical positions, and heat and dust out of the cabinets, with the progressive exceptions above put on a plan rather than left to chance.

Run a second test: ask the technicians who fix each recurring failure what they notice first. "It gets noisy about a week before" is a P-F interval and a sensor specification in one sentence. "It just stops" is a weaker answer, because technicians say it about failures that were plainly visible in trend data nobody was looking at, so treat it as a reason to look outside human sensing range rather than as a verdict.

What vibration, current, temperature, oil and cycle-time data each detect

Vibration is the richest signal for rotating machinery because its frequencies are diagnostic, not merely alarming. Defect frequencies follow from bearing geometry and shaft speed, so a spectrum says which element is damaged, though cage slip puts the real lines a per cent or two off the calculated value. Imbalance shows at running speed, misalignment typically at twice, gear problems at mesh frequency with sidebands, and envelope techniques pull early bearing impacts out of the high-frequency region before the overall level moves. The caveats matter as much: mounting sets the usable frequency range, variable-speed machines need order tracking, low shaft speeds are difficult, and vibration is near blind to electrical and process faults.

With no baseline of your own, the evaluation zones in ISO 20816-3 give a defensible starting threshold for overall velocity on general industrial machines inside the power and speed range that part covers. They are an overall severity judgement in a band running from roughly 10 Hz to 1 kHz, which is the region an early bearing defect moves last, not a bearing defect threshold. Early bearing work needs its own baseline or a technique-specific limit.

Motor current signature analysis is measured at the panel rather than on the machine, so a reading needs no access to the driven equipment and no production stop. Installing it is another matter: permanent current transformers go inside an MCC bucket or drive cabinet, which is work inside the arc flash boundary by a qualified person, usually under permit and usually with an isolation. Budget it as electrical work, not as a sensor.

It sees rotor bar breakage and eccentricity as sidebands around line frequency, and because current reflects load it catches gross mechanical change downstream: a jamming conveyor, a blocked impeller, a worn tool drawing more torque. It is coarser than vibration for bearings, and on an inverter-fed motor the classical method struggles, because the sidebands sit relative to a supply frequency that is now moving and the drive's switching harmonics sit on top of the spectrum. It wants steady-speed, steady-load capture and often underperforms even then, so where most motors are inverter-fed, the drive's own current, torque and DC bus data is the better source and the cheaper one.

Temperature is cheap, reliable and late: a hot bearing is already damaged. Where it arrives early is electrical: a thermographic survey of panels and terminations finds loose connections before they burn. Ultrasound is the earliest practical indicator of bearing lubrication problems and the standard tool for air leaks and arcing.

Oil analysis reports wear particle count and morphology, viscosity, water ingress and additive depletion, and because the particles identify which metal is wearing they often identify the component. It suits gearboxes, hydraulics and compressors, and is useless for sealed grease-lubricated bearings. Sampling discipline — same point, same conditions, same interval — decides the programme's value more than the laboratory.

Cycle-time and process drift is the most underrated signal, because the machine is already producing it: cycle time creeping up, a servo following error growing, hydraulic pressure decaying, a torque curve changing shape, a heater's duty cycle climbing to hold the same setpoint. These see fouling, wear and slippage no bolt-on sensor sees, because they measure the machine doing its job rather than a proxy for its health. Collecting them costs far less than new instrumentation but not nothing: it needs somewhere to retain the data beyond the HMI's ring buffer, tag access a machine builder may treat as a commercial conversation, and scan rate and deadband set so the drift is not compressed flat on the way into storage.

How much failure history does a predictive maintenance model need

A physics or rules approach needs no failure history at all. Defect frequencies from geometry, severity bands from a standard, a pressure differential limit across a filter, a temperature rise above the machine's own normal for the same load: all of it encodes engineering knowledge instead of demanding data. That is why a first deployment usually does not need machine learning at all, and why a proposal that opens with a model is answering a question you have not yet asked.

Anomaly detection learns the normal operating envelope from healthy running and flags deviation. It needs no labelled failures, but it needs enough healthy data to cover the legitimate variation of product, speed, material lot and shift. It tells you something changed, not what changed, and will flag a changeover as enthusiastically as a failing bearing unless the context of order, recipe, speed and tool is attached to the measurement.

Remaining-useful-life estimation needs run-to-failure examples of the same mode on the same class of asset, in numbers very few plants possess: a machine that fails twice a year produces two examples a year, possibly of two different modes. The way around the shortage is breadth instead of depth, and a fleet of nominally identical pumps measured today gives a population to compare against with no history at all. The outlier is a candidate rather than a fault, though: foundation stiffness, pipe strain, suction conditions, duty point and time since the last rebuild produce two- or threefold spreads between healthy units, so a first pass returns machines to inspect and installation problems to fix. If your plant has one of everything, that option is closed.

How to start with no usable failure history

Most plants begin with almost none, because past failures were recorded as a downtime reason and a line of free text. Design the first year to fix that rather than pretend otherwise. Baseline before you alert: instrument, collect, and stay silent through enough production variety to know what normal looks like. Alerting before normal is established is how a project buries its technicians in alarms in its first month and never recovers.

Hold any model to the thresholds you already have and make it earn its place by beating them. Then make every alert produce a label: someone inspects and records what was found, as degradation confirmed, a different mode confirmed, nothing found, or a process cause. That is the training data you did not have, and it accumulates only if the closing step is a coded field.

False alarm rates and why the maintenance team stops trusting alerts

Behind every alert is a threshold trading missed failures against false alarms, and no setting removes both. The two errors are asymmetric, but not in the way the slogan suggests. What a miss costs varies enormously by asset: on a spare pump it is an inconvenience, on the constraint it is the day's output plus secondary damage and the scrap in the machine, and on anything with a safety consequence it is not on this table at all. What a run of false alarms costs is fixed and total, because if technicians stop opening alerts, nothing downstream of the alert exists. So tune per asset against that asset's own cost of a miss, inside a false alarm budget that applies to the whole system.

Base rates make this harder than intuition suggests. Under continuous monitoring the number of false alerts is the false positive rate multiplied by how often the detector evaluates, which may be every minute on every asset, while true alerts are capped by how often the asset actually fails, a handful of times a year at most. A detector evaluating that often needs an extraordinarily low false positive rate before true alerts outnumber false ones, and accuracy says nothing about it, since a detector that never fires is accurate almost all of the time. So the first question to a vendor is not accuracy but how many alerts per week per asset the system produced elsewhere, and what fraction were confirmed on inspection.

Then budget the alerts: decide how many a technician can investigate in a week without displacing planned work, and tune to that number. Send evidence with the alert — the trend, the frequency band that moved, the sister machine for comparison — and allow three outcomes rather than two: act now, watch with a re-measurement date, or dismiss with a recorded reason.

Predictive maintenance ROI: downtime cost and the OEE arithmetic

Most business cases overstate the benefit by valuing all downtime at lost revenue, and get found out in year two when the savings appear in no account.

The correct starting point is narrower: the value is the difference between an unplanned stop and the planned intervention replacing it, multiplied by the number of events you actually convert. An unplanned failure costs the part, the overtime labour, the secondary damage — a bearing that runs to destruction often takes the shaft and housing with it — the product scrapped inside the machine, expedited freight, and the disruption to everything scheduled behind it. The planned version costs the part and the labour in a window you chose. That gap is the prize, and it is not the same as the full cost of the downtime.

Whether the lost minutes carry revenue value depends on whether the asset is the constraint and the plant is capacity-limited. If the line is the bottleneck and you sell everything you make, an hour lost is margin lost for good. If there is slack, the order is caught up later in the week, and the honest cost is the overtime premium, not the revenue.

OEE is where the losses become visible. An unplanned stop hits availability directly, but also performance through the slow ramp after a restart and quality through startup scrap, so an availability-only calculation undercounts it. What OEE cannot do is say what a lost hour is worth; only the constraint analysis does. Then subtract the recurring human time to triage alerts and maintain models as machines and products change — the line omitted from every proposal, and the one that decides whether the system still runs in year three.

Spare parts planning: warning time versus supplier lead time

A prediction is useful only if you can act inside the warning it gives, which makes this a comparison of two numbers for every monitored mode: the lead time of the signal, and the lead time of the part. Four weeks of warning on a part you stock is a scheduling decision. Four weeks on a gearbox with a five-month lead time is not a prediction, it is a countdown. So the parts side belongs in the first phase: whoever triages an alert should see whether the part is on the shelf, held on another line, or not owned at all, and the work order should reserve it when raised, not when the technician walks to the store and finds the bin empty.

Unknown failure timing is only one of the reasons safety stock exists. Supplier lead time variability, minimum order quantities, freight and obsolescence risk on older machines, and the fact that one part number usually covers several assets and several modes you are not monitoring, are all still there after the signal arrives. A warning that reliably exceeds the supplier lead time removes one reason on one asset, which is rarely enough to move a stock level by itself. Convert warnings into planned windows first and revisit stock levels after a year of evidence.

Why the CMMS work order loop has to close

The chain from signal to changed outcome has more links than the diagram in any proposal. A measurement crosses a threshold. Someone is notified, and that someone has a name and a shift, not a distribution list. They triage against the evidence and choose to act, watch or dismiss. If they act, a work order is raised with the suspected mode populated, the part reserved, and a window requested. Planning grants that window, which means the production schedule has to receive a maintenance request as a constraint rather than an argument in a corridor. Then the step everything depends on: the technician records what was found, in a coded field, against the alert that sent them.

Break any link and the rest is decoration. An alert landing on a dashboard nobody opens. A work order raised in a system that knows nothing about spare parts. A window negotiated verbally every time. And most damaging, a completion step that captures "replaced bearing" as free text, because that sentence cannot be counted, compared across assets, or used to train anything.

A structured failure coding scheme, agreed once and used every time, is the difference between a maintenance history and a pile of notes. It is also why a predictive system cannot replace a CMMS: it produces a trigger, and the CMMS turns that into work, parts, a record, and the numbers that say whether any of it worked.

What a realistic first predictive maintenance implementation looks like

Five to fifteen assets, not a plant. They come out of the twelve-month classification above: where the money sits in progressive mechanical degradation, where the asset constrains output or the failure damages product, and where a technician can already describe the warning sign.

Begin with data you already own. Drive currents, cycle times, pressures and temperatures the PLC already produces cost far less to collect than new instrumentation and often answer the question outright, given a historian to retain them and tag access to read them. Add instrumentation only where existing signals cannot see the mode: permanent accelerometers where justified, route-based portable collection elsewhere, oil sampling on the sumps, and a thermographic survey of the electrical panels at representative load, through fitted infrared windows or under a live-work permit, because a survey run at a tenth of normal load finds nothing and gets filed as a clean result.

Write acceptance criteria before the sensors arrive, in terms someone will still understand a year later: alerts per week per asset, the fraction confirmed on inspection, the warning lead time achieved, and the number of unplanned stops converted into planned windows. Not model accuracy, which nobody on the floor can act on.

Name an owner on the maintenance side, not the IT side, and review at six months against those leading numbers, because they are the only ones that will have moved. The lagging measure, unplanned stops avoided, needs enough events to be more than noise, and on ten assets that fail a few times a year, only some of them in the way the system watches for, that is well past a year. Judging at six months on unplanned stops is how a working system gets cancelled and a useless one gets expanded. Go into that review willing to conclude that on some assets, condition-based rounds with an instrument and a trained person are the right answer. That is a successful project, not a failed one.

Where the maintenance module fits

Meta Smart Factory's Maintenance Management System module covers the CMMS side of the loop above: asset hierarchy and history, preventive plans on time or counters, work orders with coded findings, spare parts reserved against the order, and MTBF, MTTR and cost-per-asset figures that fall out of those records rather than being reconstructed afterwards. Downtime captured in MES raises the notification without anyone re-typing it, the requested window reaches APS as a scheduling constraint instead of a phone call, condition data arrives over the same IIoT and OPC UA connectivity, and trending and anomaly detection run in the AI and Machine Learning module once there is a baseline worth using.

The sequence matters more than the module list. Classifying twelve months of your own failures costs nothing but time, decides whether sensors are the right purchase at all, and is a more useful first conversation than watching an alert fire on someone else's bearing.

Discuss This With Our Experts