Three ways to intervene, with the right names
On the shop floor people say “run to failure,” “on a schedule,” “on condition.” UNI EN 13306:2018, the European maintenance terminology standard, uses other names, the ones you find in specifications and audits. The definitions below are restated in our own words; the full text is available from UNI.
| What you call it | Standardized term | When the intervention is triggered |
|---|---|---|
| Run to failure | Corrective maintenance | When the machine has already stopped working |
| On a schedule, by hours | Predetermined preventive maintenance | When hours or months decided at a desk expire, whatever the condition |
| On condition, CBM | Condition-based maintenance | When a measurement says the condition has changed enough |
| Predictive, PdM | Predictive maintenance | As above, but with a date estimated from the trend of the measurements |
Predictive is not a fourth category: in the structure of the standard it is a subset of condition-based maintenance. And predetermined and condition-based are both branches of preventive maintenance, so “preventive” is not a synonym for “scheduled.” If you measure and intervene when a threshold is exceeded, you are already doing condition-based maintenance. To call it predictive you need a forecast built on the trend of the degradation parameters, and in any case no measurement gives you the date of the failure. “Ordinary” and “extraordinary” are not strategies: they sit on the economic axis of UNI 11063:2017, an Italian standard on maintenance classification.
Scheduled maintenance is not a mistake
Whoever replaces bearings and belts at fixed intervals is not doing it wrong: in many cases it is the economically correct choice. Peer-reviewed literature says so too: the advantage of condition-based maintenance over fixed-interval maintenance is not universal; it depends on how the component degrades, on how severe the failure is, and on how accurate the measurement is (de Jonge, Teunter, and Tinga, 2017). The cases where the schedule wins are precise.
The reverse is also true. Opening up a machine that is running fine is not free. The 1978 Nowlan and Heap report, on United Airlines civil aviation components with data from the 1960s and 70s, observes that on complex assemblies scheduled overhaul can raise the overall failure rate, introducing infant mortality into a stable system. It is not a study of industrial machinery and does not transfer as is, but you recognize the mechanism: after certain reassemblies, failures appear that weren’t there before. It is not an argument against the schedule; it is an argument for not applying it to everything.
The schedule is the right choice when:
- The component has a recognizable, repeatable wear life: filters, oil, gaskets, belts.
- The intervention is required by law or requested by the manufacturer for the warranty.
- The failure mode has no symptom measurable in advance: ISO 17359 anticipates the case and refers to other strategies.
- The intervention is low-cost within an already scheduled shutdown: measuring would cost more than the spare part.
- The time between first symptom and failure is shorter than the time you need to react.
The math first, then the sensor
ISO 17359:2018 gives guidelines on how a monitoring program is set up. It is not binding and it is not certifiable: nobody is “ISO 17359 compliant.” The useful part is the order of the steps, which we summarize in our own words without reproducing the standard’s diagram.
The point that surprises almost everyone: the economic analysis is the first block, not the last. In its commentary the standard refers to return on investment and life cycle cost. The question “how much does it earn me” comes before “which sensor do I buy,” and it’s the standard saying so. Thresholds are not a fixed given: the review cycle puts them back on the table.
- Cost-benefit analysis (Clause 5): checking whether monitoring pays for itself, counting what the machine costs over its life and what is lost when it stops.
- Inventory of the machines, their function, and their operating conditions (Clause 6).
- Criticality assessment across all machines: out of it come the priority order and the list of those to leave out (Clause 7.2).
- Failure mode analysis: which parameter can signal them, and whether it is measurable (Clauses 7.3 and 7.4).
- Method, measurement points, interval, initial thresholds and baseline, then trending, action, and periodic review (Clauses 8-11).
The cost of downtime, item by item
No standard gives you a formula for the cost of downtime: ISO 17359 cites it as an item to consider, UNI EN 15341:2022 gives maintenance performance indicators, not a cost model. The calculation is yours to do: copy the table and fill in the last column by hand.
| Item | How to estimate it | Example | Your value |
|---|---|---|---|
| Duration of the downtime (basis, not added) | From the call to restart at full rate | 9 h | — |
| In-house labor and overtime | Man-hours of restoration and recovery outside normal hours | €210 | — |
| Emergency external service | Travel, out-of-contract rate, spare part | €800 | — |
| Production lost or shifted | Parts not made, or moved to another machine | €480 | — |
| Delivery delay | Penalties, urgent transport | €300 | — |
| Consequential damage | Workpiece lost, downstream damage | €200 | — |
The example is constructed by us, not market data: a total of €1,990 for nine hours of downtime on a machining center. If you work to order, you often don’t lose the margin: you shift it. You recover on a Saturday or on another machine, and the revenue comes in anyway. But two rows remain true, the overtime and the delivery delay: fill them in honestly, otherwise the math tells you zero, and it isn’t zero.
The mistake to avoid: the benefit is not the whole downtime
Careful. The benefit of monitoring is not the full cost of the downtime avoided. A failure seen coming still costs: the bearing gets replaced anyway, the machine stops, the technician gets paid. What changes is when and how.
The correct formula is a difference: benefit per event = cost of the unplanned intervention minus cost of the same intervention planned. Then you multiply by the events per year you actually expect to catch, which on a single machine is often a fraction of an event. On the table’s example: labor from €210 to €120 with no overtime, external service from €800 to €400 at the ordinary rate, delivery delay from €300 to zero, consequential damage from €200 to €50, lost production keeps the full €480. The same failure, seen coming, costs €1,050: the benefit is €940 out of €1,990. Whoever presents the whole downtime as savings is inflating your numbers.
When the answer is no, and where to start
On some machines the honest answer is no: a redundant machine with its twin ready to go, a spare on the shelf and a one-hour replacement, downtime that blocks nothing downstream, or a failure mode with no measurable symptom. The list of machines to leave out is a result of the method, not a shortcoming of it.
| Factor | Weight | How to score from 1 to 5 |
|---|---|---|
| Cost of downtime and lost production | x3 | 1 = negligible, 5 = stops the line or the customer |
| No redundancy and no spare at hand | x3 | 1 = twin and part ready, 5 = one-of-a-kind machine, months of lead time |
| Failure frequency and repair time | x2 | 1 = never, fixed in an hour, 5 = recurring, days of work |
| Consequential damage, safety, environment | x2 | 1 = none, 5 = risk to people or a spill |
| Feasibility of the measurement | x1 | 1 = inaccessible point or variable regime, 5 = easy access, repeatable cycle |
Weighted sum, maximum 55. Rank the machines from the highest score and start with the first two or three, not with twenty. An honest distinction: assessing criticality with a score on weighted factors is suggested by ISO 17359 (Clause 7.2), and nearly all of the grid’s factors come from there; the feasibility of the measurement comes from the clause on the monitoring method. The weights and the 1-to-5 scale are ours: calibrate them on your own plant.
Before buying anything
- I have written down the cost of downtime on this machine, not a sector average.
- I have calculated the difference between an unplanned and a planned intervention, not the whole downtime.
- I know which failure mode I want to catch and which parameter shows it.
- I know, as an order of magnitude, how much time passes between first symptom and failure.
- I know who looks at the data and what they do when an alarm comes in.
Download the PDF
The same guide in printable format · PDF, 556 KB
Sources and disclaimer
UNI EN 13306:2018 Maintenance, maintenance terminology (definitions restated, full text available from UNI, the Italian standards body). UNI 11063:2017 Ordinary and extraordinary maintenance. ISO 17359:2018 Condition monitoring and diagnostics of machines, general guidelines, Clauses 5, 7.2, and 8.3. UNI EN 15341:2022 Maintenance, maintenance key performance indicators (KPIs). Nowlan F.S., Heap H.F., Reliability-Centered Maintenance, United Airlines for the Office of the Assistant Secretary of Defense, December 29, 1978, DTIC report AD-A066579, unlimited distribution. de Jonge B., Teunter R., Tinga T., Reliability Engineering & System Safety 158 (2017), pp. 21-30. The numerical examples are constructed by us as examples; they are not market data. Informational document. The cited standards are available from UNI: what is reported here is a restatement, not the normative text.