How to Choose a Maintenance Strategy for Each Asset: Run-to-Failure, Time-Based and Condition-Based Maintenance, and How Criticality Decides Between Them

Not Every Machine Deserves the Same Care

Published 2026-09-07 · EuroQuest International

Quick summary

  • There is no best strategy, only a best fit per asset. Run-to-failure, time-based and condition-based maintenance each win on a different class of equipment, and applying one policy to the whole plant guarantees waste at one end and surprises at the other.
  • Criticality decides, not age or cost. The question is what happens when the asset stops: who is exposed, what production is lost, what is released, and how long recovery takes.
  • Run-to-failure is a legitimate choice, not neglect. It is the right answer for low-consequence, cheap, quickly replaced items, and writing it down is what separates a decision from an oversight.
  • Time-based intervals need a failure pattern to justify them. Where failures are random rather than wear-related, a calendar interval adds cost and introduces fresh assembly errors without reducing risk.
  • The strategy drifts back unless someone owns it. Deferred tasks, unresolved repeat failures and blame directed at operators are the three signals that the written policy and the real one have parted company.

Most plants do not choose a maintenance strategy. They inherit one. A shift supervisor who left in 2014 set the greasing intervals, a vendor manual set the overhaul cycle, and everything the manual did not mention gets fixed when it breaks. The result is usually the same in every industry: a handful of assets are maintained far more often than their failure behavior justifies, a different handful are quietly running to destruction, and nobody has written down why either group is treated the way it is.

Choosing deliberately is not complicated, but it does require answering questions in a particular order. This guide sets out the strategies available, how to sort assets between them, what data you need before changing anything, and how to tell whether the choice is holding. It is about the decision itself. It is not a job description for the people who carry it out, and it is not the after-the-fact analysis you run once something has already failed, which is a separate discipline covered in the guide to conducting a root cause analysis.

On this page

  1. What are the maintenance strategies to choose between?
  2. How do you decide which asset gets which strategy?
  3. When is run-to-failure the right answer?
  4. What data do you need before changing anything?
  5. Why do maintenance strategies drift back?
  6. How do you know the strategy is working?
  7. Frequently asked questions
76.0%
Capacity utilization for US manufacturing in July 2026, some 2.2 percentage points below its 1972 to 2025 average, on Federal Reserve figures
2.4%
Rise in US manufacturing labor productivity in the second quarter of 2026, as output rose 5.4 percent and hours worked rose 2.9 percent, per the Bureau of Labor Statistics
$4.5bn
Property damage across the 94 serious chemical incidents in 31 states covered by the four volumes of CSB incident reports, alongside 16 fatalities

What Are the Maintenance Strategies to Choose Between?

Strip away the vendor vocabulary and there are three real choices, plus a fourth that is a refinement of the third rather than a category of its own. Each is defined by what triggers the work.

Run-to-Failure: the Trigger Is the Failure

Nothing is scheduled. The asset is operated until it stops, and then it is repaired or replaced. This is the cheapest policy to administer and the only one that extracts the full usable life of a component. Its cost is concentrated entirely in the consequence of the failure, which is why it is defensible on a corridor light fitting and indefensible on a pump whose loss shuts a line. Chosen deliberately, with spares on the shelf, it is sound engineering. Arrived at by default, it is simply an absence of a decision.

Time-Based or Usage-Based: the Trigger Is the Calendar or the Counter

Work is performed at a fixed interval: every 500 running hours, every quarter, every 10,000 cycles. This is the regime most plants default to because manuals are written in its language and because it is easy to plan and audit. It works when a component genuinely wears out, meaning its failure probability rises with age in a way the interval can be set against. It works poorly when failures are random, because replacing a healthy part on a schedule cannot lower a risk that does not increase with time, and each intervention introduces a fresh chance of installation error.

Condition-Based: the Trigger Is a Measurement

The asset is monitored, and work is triggered when an indicator crosses a threshold: vibration, temperature, oil particle count, motor current signature, ultrasound. Intervention happens when the equipment says it is needed rather than when the calendar does. The tradeoff is that the monitoring itself costs money and, more importantly, attention. A vibration route walked religiously and never analyzed is a cost with no benefit, which is why this regime lives or dies on whether someone competent reads the results. Building that reading capability is the substance of maintenance planning and reliability engineering work rather than a software purchase.

Predictive Analytics: a Refinement, Not a Fourth Category

Adding models to condition data to estimate remaining useful life is a real advance, but it is still condition-based maintenance with better inference. It inherits every weakness of the underlying data, and it cannot rescue an asset register that is wrong about what equipment exists. Plants that buy the analytics before fixing the register spend a year discovering that the sensors are attached to assets nobody can identify.

How Do You Decide Which Asset Gets Which Strategy?

The sorting question is not how expensive the asset is or how old it is. It is what happens when it stops. Two identical motors can warrant opposite policies if one drives a cooling water pump with no standby and the other drives a spare conveyor. Consequence, not equipment type, is the sorting variable, and it is the reason asset management and equipment reliability is treated as a business decision rather than a technical one.

A workable criticality assessment asks four things about each asset: whether its failure can injure someone or release something, whether production stops or merely slows, whether a standby exists and how quickly it can be brought in, and how long a replacement takes to obtain. Assets that score high on safety or environmental consequence go to condition-based monitoring regardless of cost, because the loss being avoided is not measured in repair hours.

Strategy Best suited to What it costs How it fails
Run-to-failure Low-consequence, low-cost, quickly replaced items with spares held Almost nothing to administer; the full consequence of the failure Applied by default to something critical because nobody assessed it
Time-based or usage-based Components that genuinely wear, and statutory or insurance-driven inspections Planned labor and parts, plus the production time the intervention takes Intervals set by habit; healthy parts replaced; errors introduced during the work
Condition-based Rotating and high-consequence equipment where a measurable indicator degrades Instruments, routes, and the analytical attention to read them Data collected and never analyzed; thresholds nobody trusts, so alarms are ignored
Redesign or eliminate Assets that keep failing the same way whatever regime is applied Engineering time and capital, once Never considered, because maintenance owns the problem and engineering owns the fix

The fourth row is the one most often missing from the discussion. An asset that fails repeatedly in the same mode under every regime is not a maintenance problem at all. It is a design or application problem being paid for out of the maintenance budget, month after month, by people who have no authority to fix it.

Key terms, used precisely

  • Failure mode. The specific way an item stops performing its function: a bearing seizing is a different mode from the same bearing running hot, and each may call for a different regime.
  • Criticality. A ranking of consequence, not of value. It combines safety and environmental exposure, production loss, redundancy, and recovery time.
  • P-F interval. The time between the first detectable sign of a developing failure and the failure itself. It sets how often you have to look, and if it is shorter than your inspection route, monitoring will miss it.
  • Mean time between failures. An average, which means it says nothing useful about a population of two identical machines. Treat it as a trend indicator, never as a prediction for one asset.

When Is Run-to-Failure the Right Answer?

More often than most maintenance policies admit. Four conditions have to hold together. The failure must not hurt anyone or release anything. Production must either continue or degrade gracefully rather than stop. A replacement must be available quickly, which usually means it is on the shelf. And the failure must not damage anything else on its way out, because a cheap coupling that destroys a shaft when it goes is not a cheap failure.

Where those hold, scheduled attention is money spent to avoid a consequence that was never expensive. The discipline is in writing the choice down. A run-to-failure decision recorded against an asset, with the spare part policy attached, is a defensible position in an audit and a useful instruction to a technician. The same asset with no entry at all looks identical on the floor and indefensible on paper, and that distinction matters most in buildings and support systems, where facility and plant management teams carry hundreds of low-consequence items that nobody has ever formally classified.

What Data Do You Need Before Changing Anything?

Less than vendors suggest and more than most plants have. Three things are non-negotiable, and none of them requires a new system.

An Asset Register That Matches the Floor

Every reliability initiative that fails early fails here. If the register lists equipment that was removed in 2019, omits the two units installed last year, and identifies four pumps by three different tag conventions, then every downstream calculation is fiction. Walking the plant with the register in hand is unglamorous, takes a week, and is the highest-return week in the entire program.

Failure History, Coded by Mode Rather Than by Fix

Most work-order histories record what was done, not what went wrong. "Replaced bearing" tells you nothing about whether the bearing was starved of lubricant, misaligned, overloaded, or simply old, and those four causes point to four different strategies. Adding a short, closed list of failure modes to the work-order close-out is a small change that becomes useful after about six months and decisive after two years. Teams that measure this properly usually find the exercise sits closer to industrial engineering and productivity improvement than to maintenance craft.

The Real Cost of an Hour of Downtime

Not the maintenance cost, which is small and visible, but the production cost, which is large and usually contested between operations and finance. Without an agreed figure, every argument about whether to monitor an asset becomes a matter of opinion, and the opinion of whoever is more senior wins. With one, the comparison is arithmetic. The wider picture is worth keeping in view: US manufacturing ran at 76.0 percent of capacity in July 2026 on Federal Reserve figures, and output per hour in the sector rose 2.4 percent in the second quarter of that year as output grew faster than hours worked. Neither number is caused by maintenance policy, but both describe the installed capacity that unplanned stoppages take out of service.

Why Do Maintenance Strategies Drift Back?

Because a strategy is a set of promises about future effort, and future effort is the first thing borrowed against when a plant is busy. Drift is rarely announced. It shows up as tasks deferred to the next window, then to the window after that; as condition routes walked but not reviewed; and as the same failure appearing three times without anyone asking why the regime did not catch it.

The most telling signal is where the explanation lands. When repeat failures are attributed to the operator, the strategy has effectively been abandoned without anyone saying so. The point is made bluntly in the US Occupational Safety and Health Administration's guide for employers on incident investigation, which states that if an investigation is focused on finding fault, it will always stop short of discovering the root causes, and that root causes generally reflect management, design, planning, organizational or operational failings rather than individual carelessness. A maintenance regime that does not fit the asset is exactly such a failing, and it will keep producing incidents that look like human error from a distance.

The consequences at the severe end are documented rather than theoretical. The four volumes of incident reports published by the US Chemical Safety and Hazard Investigation Board cover 94 serious chemical incidents across 31 states, involving 16 fatalities, 75 serious injuries and more than $4.5 billion in property damage. Those are chemical process incidents in one country, so they are not a general failure rate for industry, but they establish the scale of what sits at the far end of a mechanical integrity decision. Keeping a strategy alive between crises is a cultural problem more than a technical one, which is why it is usually addressed as part of building a culture of continuous improvement.

How Do You Know the Strategy Is Working?

Not by counting completed work orders, which rise whether or not anything improved, and not by maintenance spend, which can fall for a year simply because work was deferred. Four measures actually move when the choice was right, and all four need a baseline taken before anything changes.

The Four That Matter

The ratio of planned to unplanned work is the headline: a regime that fits its assets converts surprises into scheduled jobs, and that ratio should move steadily in one direction. Repeat failures on the same asset in the same mode should fall to near zero, and any that persist point at the redesign row of the table rather than at the technician. Time lost to unplanned stoppages, valued at the agreed hourly figure, is the number finance will accept. And the proportion of condition alarms that led to an intervention tells you whether the thresholds are trusted, because a monitoring system whose alarms are routinely dismissed has already stopped working. Assembling those four into something a plant reviews monthly is the practical end of operational excellence and process optimization.

Review on a Fixed Date, Not After an Incident

A strategy reviewed only after something goes wrong will always be revised in the direction of the last event, which is how plants end up monitoring one pump exhaustively while an identical one three meters away is untouched. An annual review on a fixed date, covering the critical assets in order, keeps the policy proportionate. It also creates the moment at which a deferred task becomes visible as a decision rather than disappearing as a habit.

Before you sign off a maintenance strategy

  1. The asset register has been walked against the plant, and tags are unique and consistent.
  2. Every asset carries a criticality rating based on consequence, and the rating has a named owner.
  3. Run-to-failure decisions are written down, with the spares policy attached to each one.
  4. Every time-based interval can be justified by a wear mechanism or a legal requirement, and any that cannot is on a list to be challenged.
  5. Each condition-monitoring route has a named person who reviews the results and a threshold that person will defend.
  6. Assets with three or more repeat failures in the same mode are flagged for engineering, not rescheduled.
  7. A baseline for the four measures exists and is dated before the first change is made.
  8. A review date is in the calendar, owned by a person rather than a department.

A useful test before committing: for any asset you are about to put on a fixed interval, ask what wears out. If nobody in the room can name the mechanism, the interval is a guess, and a guess that costs production time every cycle is worse than an honest run-to-failure decision.

Where Teams Build This Capability

Maintenance strategy is decided badly when it is decided by one function alone, because operations owns the consequence, maintenance owns the effort, and finance owns the number that settles the argument. Practitioners take this work in Paris, Istanbul, Kuala Lumpur, Manama and Barcelona, and the wider field sits under engineering and operations management.

Frequently Asked Questions

What are the main types of maintenance strategy?

Three, distinguished by what triggers the work. Run-to-failure waits for the breakdown. Time-based or usage-based maintenance acts on a calendar or a counter. Condition-based maintenance acts when a measured indicator crosses a threshold. Predictive analytics is often listed as a fourth, but it is condition-based maintenance with better inference rather than a separate category, and it inherits the quality of the underlying data. A fifth option deserves a place in the discussion: redesigning or removing an asset that keeps failing the same way whatever regime is applied to it.

Is run-to-failure ever acceptable?

Yes, and it is the correct choice more often than most written policies admit. Four conditions have to hold together: the failure cannot injure anyone or release anything, production continues or degrades gracefully rather than stopping, a replacement is available quickly, and the failure does not damage anything else on its way out. What separates a legitimate run-to-failure decision from neglect is that it is written down against the asset, with the spare part policy attached, so a technician and an auditor both see a decision rather than an omission.

How do you set a preventive maintenance interval?

Start from a wear mechanism, not from the manual alone. A fixed interval only reduces risk where failure probability rises with age or use, so the first question is what physically wears out and how quickly. Where the answer is known, set the interval comfortably inside that life and adjust it against failure history rather than against habit. Where nobody can name a mechanism, the interval is a guess that costs production time every cycle and introduces a fresh chance of assembly error, and condition monitoring or an honest run-to-failure decision will usually serve better.

What is criticality analysis in maintenance?

It is the ranking that decides which strategy each asset gets, and it measures consequence rather than value. Four questions do most of the work: can the failure injure someone or release something, does production stop or merely slow, does a standby exist and how fast can it be brought in, and how long does a replacement take to obtain. Two identical machines can end up with opposite policies on this basis. Anything scoring high on safety or environmental consequence goes to condition-based monitoring regardless of its purchase price, because the loss being avoided is not measured in repair hours.

How is choosing a maintenance strategy different from root cause analysis?

They sit on opposite sides of the failure. Root cause analysis is retrospective and single-event: something has already broken, and the work is to establish why so the same conditions do not recur. Choosing a strategy is prospective and population-wide: it decides, before anything breaks, which regime each asset should be under and what evidence would justify changing that. The two connect in one direction. A pattern of root cause findings is among the best evidence available for revising a strategy, which is why coding failure modes at work-order close-out pays off long after the individual investigations are forgotten.

Give every asset a regime someone can justify

EuroQuest International runs practitioner training in maintenance planning and reliability engineering, asset management, and operational excellence across Europe, the Gulf and Asia.

Explore engineering and operations programs