Skip to content

Free tool

Maintenance strategy selector

To choose how to maintain a failure mode, ask three questions about its consequences: would anyone notice it, could it hurt someone or breach an environmental limit, and does it stop output. Then try the tasks in order: a condition check if there is a warning sign that lasts longer than the time you need to act, a time-based replacement only if failures rise at a known age, and otherwise the default for the consequence: redesign where someone could be hurt, run to failure where it only costs money, a failure-finding test where the failure is hidden. That is the order of the RCM decision diagram Nowlan and Heap published in 1978. A bearing with 6 weeks of vibration warning is checked every 3 weeks; a relief valve with a 20-year MTBF that must be 99.5% available is tested every 73 days.

Check intervalequalsthe smaller of P-F interval divided by 2 and (P-F interval − time to act)

Failure-finding intervalequals2 × (1 − A) × MTBF

Failure-finding intervalequals2 × MTBF × MTED divided by MMF

The P-F interval is the warning time, from the point P where a coming failure can first be detected to the failure F; the check interval rule is from the ABS Guidance Notes on Reliability-Centered Maintenance (section 4, 4.1). A is the availability you need from a protective device, MTBF its mean time between failures, MTED the mean time between demands on it and MMF the mean time between multiple failures you can tolerate, all in the same unit (ABS, section 5, equation 5). The failure-finding formulas assume random failures and an interval under a tenth of the MTBF.

One failure mode at a time: a specific way the asset fails, such as "drive-end bearing wears", not "the compressor fails". The questions follow the order of the RCM decision diagram. "Not sure" takes the cautious default and marks the answer provisional.

Evident in normal work?

Would the people running it notice this failure on their own, during normal work?

What counts

Yes if it stops the machine, sets off an alarm, shows on a gauge someone watches or changes the output. No for a protective device or a standby unit that can sit failed without anyone knowing, such as a relief valve stuck shut or a backup pump that will not start.

Safety or environment?

Could this failure hurt someone, or breach an environmental or legal limit?

What counts

Count injuries, fires, releases and permit or legal breaches, not cost. Include damage the failure does to other parts that could then hurt someone.

Affects output or quality?

Does it stop or slow output, or hurt quality or delivery?

What counts

Yes if the line stops, runs slower, makes scrap or misses shipments while it is failed. No if the repair can wait for a convenient time and the only cost is the repair itself.

Warning sign before failure?

Is there a warning sign you can detect before it fails?

What counts

A potential failure: a condition that says the failure is on its way, such as rising vibration or temperature, a crack, wear depth, metal in the oil, a leak or a rising pressure drop across a filter. It has to show before the failure, and you need a check that can find it.

P-F interval and time to act

How long does the warning last, and how long do you need to act?

What counts

The P-F interval runs from the point P where the sign can first be detected to the point F where the failure happens. Use the shortest warning you have seen or would expect. Technicians, past condition readings and the maker are the usual sources.

Plan the job, get the parts, schedule the stop.

Warning time consistent?

Is the warning time roughly the same from one failure to the next?

What counts

Nowlan and Heap make a reasonably consistent time between the potential failure and the failure a condition for a condition task. If it ranges from days to months, a check interval cannot be set with any confidence.

Check practical?

Can you check every 3 weeks in practice, with a method that finds the sign?

What counts

The check has to find the sign reliably (the right instrument and someone trained to read it) and be possible at that frequency without more stops than you can accept.

Checks cheaper than failures?

Over a year, would the checks cost less than the failures they prevent?

What counts

Compare the cost of the checks with the repair cost plus the cost of lost output for the failures they would catch.

Recommended strategy

Condition-based

Check for the warning sign at a fixed interval and plan the repair when the sign is found.

Consequence

Operational

The failure stops or slows output, or harms quality or delivery.

Interval basis

Check every 3 weeks

Half the P-F interval of 6 weeks; leaves 3 weeks to act, 1 week needed.

Why: the path through the questions

  1. 1Evident in normal work? YesThe failure is evident, so its own consequences decide what is worth doing.
  2. 2Safety or environment? NoNo safety or environmental consequence, so the choice is an economic one.
  3. 3Affects output or quality? YesOperational: a task is worth doing if it costs less than the repairs plus the lost output it prevents.
  4. 4Warning sign before failure? YesA condition check may work, if the warning lasts long enough.
  5. 5P-F interval and time to act 6 weeks warning, 1 week to actCheck every 3 weeks (half the P-F interval), which leaves at least 3 weeks to act.
  6. 6Warning time consistent? YesThe warning time can be relied on.
  7. 7Check practical? YesThe check can be done at that interval.
  8. 8Checks cheaper than failures? YesThe checks pay for themselves.

Data to collect next

  • Record each reading against its alert limit. Shorten the interval if the sign is found late; lengthen it if readings stay flat over several checks.

Supports a decision for one failure mode; it does not replace an RCM analysis with the people who run and maintain the asset. For the full study, use the RCM worksheet.

Summary table

Add each failure mode as you finish it, then copy the table into your PM plan or CMMS import sheet. The example rows are the worked example below.

AssetFailure modeStrategyInterval basisOwnerRemove
Air compressor C-2Motor drive-end bearing wearsCondition-based maintenanceCheck every 3 weeksReliability technician
Air compressor C-2Pressure relief valve sticks shutFailure-finding testTest every 0.2 years (73 days)Maintenance planner
Air compressor C-2Panel indicator lamp burns outRun to failureNo scheduled task; repair on failureOperator

Email me my results

Optional. Get your answers, the recommendation and the reasoning in your inbox, with a link back to this tool.

We may follow up about LeanSuite. See our Privacy Notice.

Next step

Want to schedule and track these tasks?

Bring this result to a 45-minute demo and we will show where it fits in LeanSuite's Professional Maintenance Tags.

How to use it

One failure mode at a time

1. Choose the asset, then its failure modes. An equipment criticality analysis picks the asset worth the time. A failure mode is a specific cause, such as "drive-end bearing wears" or "relief valve sticks shut", taken from work orders and from the technicians who fix it.

2. Answer the consequence questions. The first one, whether anyone would notice the failure in normal work, is the one most often answered wrongly: protective devices and standby units can sit failed for months without a sign.

3. Answer the task questions from evidence. Warning signs and warning times come from the technicians and from condition readings, failure ages from your work orders. "Not sure" takes Nowlan and Heap's cautious default answers: hidden, safety, operational, a condition check assumed workable, a scheduled replacement assumed not until there is data. The result is then marked provisional.

4. Read the reasoning, then set the interval. Round the interval down to a slot in your schedule you can keep, and shorten it for a failure that could hurt someone.

5. Add it to the table and do the next one. Copy the table into your PM plan. Review each interval as findings come in: a sign found late means a shorter interval; readings that stay flat over several checks allow a longer one.

Worked example

Illustrative numbers, not a benchmark

Three failure modes of a rotary screw air compressor that supplies a line. The warning time, the MTBF and the availability target are illustrative, not data from a real machine.

  1. 1Drive-end bearing wears. Evident (it trips the motor), no safety consequence here, and the line loses air: operational. Vibration analysis picks up the defect about 6 weeks before failure, and getting the bearing and a stop takes 1 week.
  2. 2Check every 6 ÷ 2 = 3 weeks. The worst case, a defect that appears just after a check, leaves 6 − 3 = 3 weeks to act, more than the 1 week needed. The checks cost less than a failed bearing and a stopped line: condition-based maintenance every 3 weeks.
  3. 3Pressure relief valve sticks shut. Nobody notices until the pressure control also fails, and that multiple failure could hurt someone: hidden, safety. No warning sign without a test and no known wear-out age, so the default is a failure-finding test, done on a test bench.
  4. 4With a valve MTBF of 20 years and 99.5% availability set by the site: 2 × (1 − 0.995) × 20 = 0.2 years = 73 days. 73 days is 1% of the MTBF, well under the one-tenth limit: test every 73 days, scheduled every 10 weeks.
  5. 5Panel indicator lamp burns out. Evident, harmless and costs only a lamp: non-operational. No warning sign, failures at random: run to failure, with spare lamps in stock.
Three rows in the summary table: condition-based every 3 weeks, failure-finding every 73 days, run to failure. These are the rows the tool opens with; change the answers to see how each one moves.

Why the questions come in this order

The tool follows the RCM decision diagram in Nowlan and Heap's report for the US Department of Defense (Exhibit 4.4). Consequences come first, because they decide what "worth doing" means:

  • Safety or environment: a task has to bring the risk to a level the site accepts, whatever it costs. If none does, redesign is required.
  • Operational: a task has to cost less than the repairs plus the lost output it prevents.
  • Non-operational: a task has to cost less than the repairs alone.
  • Hidden: a task has to give the protective function the availability you need. The default is a failure-finding test.

Within each branch a condition check is tried first, because it replaces a part only when the part needs it; Nowlan and Heap call it the most desirable task whenever it applies. Environmental and legal consequences sit with safety, as in later RCM standards such as MIL-STD-3034A and the NASA RCM guide.

The six failure patterns

United Airlines plotted the chance of failure against age for its aircraft components and found six shapes. Nowlan and Heap published them in 1978 (Exhibit 2.13), with the share of the items studied in each:

  • A. Bathtub, 4%High when new, then low and steady, then a wear-out zone.
  • B. Wear-out, 2%Steady or slowly rising, then a pronounced wear-out zone.
  • C. Slowly rising, 5%Rises gradually with age, with no clear wear-out age.
  • D. Low, then steady, 7%Low when new or just repaired, then rises quickly to a steady level.
  • E. Random, 14%The same chance of failure at every age.
  • F. Early life, 68%High when new or just repaired, then steady or very slowly rising.

In the report's words, some 89 percent of the items had no wear-out zone, so an age limit could not improve them; 11 percent (A, B and C) might benefit from one. The NASA RCM guide (2008) adds a Swedish study (1973) and a US Navy study (1982) and puts random failures at 77 to 92 percent across the three.

These are aircraft and naval items, not plant equipment, and we found no comparable published study for plants that we could check. What carries over is the test: replacing on a schedule helps only a failure mode with a clear wear-out age that most units reach. Early-life failures (pattern F) point to installation, start-up and repair quality, and the NASA guide notes that a scheduled overhaul often adds more of them.

Axes: chance of failure (vertical) against age since new or since repair. Some later texts letter the patterns in a different order; these are Nowlan and Heap's.

The P-F interval and the check interval

Nowlan and Heap define a potential failure as "an identifiable physical condition which indicates a functional failure is imminent". A condition check works only if that condition can be detected, the time from it to the failure is reasonably consistent, and there is time to act.

  • Check at no more than half the P-F interval (ABS, 4.1). With a 6-week warning, check every 3 weeks.
  • Leave time to act. In the worst case the sign appears just after a check, so the time left is the P-F interval minus the check interval. If that is shorter than the time you need to plan and do the repair, shorten the interval. If the warning is not longer than the time to act, no check works.
  • Shorten it further for higher-risk failure modes or when the P-F estimate is a guess (ABS, 4.1). Once a reading passes its alert limit, the NASA guide says to cut the monitoring interval to between a third and a quarter of what it was.
  • Learn the real P-F interval. ABS gives the example of pumps whose weekly vibration readings caught defects that then ran 6 to 8 weeks before repair without failing: the P-F interval is at least 6 weeks, so the check can move to every 3 weeks.

The failure-finding interval

A protective device that fails at random and is tested every T is failed, on average, for about T ÷ (2 × MTBF) of the time. Set that equal to the unavailability you can accept and solve for T: T = 2 × (1 − A) × MTBF. If you know how often the device is called on (MTED) and how rarely you can accept the multiple failure (MMF), the unavailability is MTED ÷ MMF.

The interval as a share of the device MTBF (ABS, Table 2)
Unavailability acceptedTest interval
0.0001 (99.99% available)0.02% of MTBF
0.001 (99.9%)0.2% of MTBF
0.01 (99%)2% of MTBF
0.05 (95%)10% of MTBF

The unavailability you accept is your site's risk rule, not a standard. ABS lists the assumptions: random failures, failure rate × interval under 0.1, test and repair times short, and the multiple failure possible only from that one demand. The test itself must not create the hazard, and should prove the whole protective function, not one part of it.

What this tool does not do

  • It does not define functions or list failure modes. An RCM analysis does that with the people who run and maintain the asset; the RCM worksheet holds the whole study.
  • It does not decide what is safe. Whether a failure could hurt someone, and what risk the site accepts, are decisions for the site and its safety professionals.
  • It does not override a law, code or insurer. Where one sets an inspection or test interval, as is common for pressure relief devices, that interval applies whatever this tool says.
  • It does not supply failure data. Every warning time, MTBF and cost is yours; the tool only shows what follows from them.
  • No claim is made that it meets SAE JA1011, the standard that sets what a process must include to be called RCM.

Sources and how this tool was checked

  • F. Stanley Nowlan and Howard F. Heap, Reliability-Centered Maintenance (US Department of Defense, 1978, report AD-A066579): potential failure (section 2.1), the six patterns (section 2.8, Exhibit 2.13), the task criteria (chapter 3), the decision diagram and default answers (Exhibits 4.4 and 4.5).
  • American Bureau of Shipping, Guidance Notes on Reliability-Centered Maintenance (2004, updated 2018): section 4 for the P-F and check intervals, section 5 for the failure-finding interval and Table 2.
  • NASA, Reliability-Centered Maintenance Guide for Facilities and Collateral Equipment (2008): chapter 3 for the four outcomes, section 4.1.5 for the three failure-pattern studies, page 4-3 for shortening the interval after an alert.
  • John Moubray, Reliability-centred Maintenance (RCM II), 2nd ed. (1997), the usual reference for the P-F interval and both interval rules, is not free to read and was not used to check this page.

Unit tests walk every combination of answers to make sure each ends in a recommendation, check the paths above against the decision diagram and the default answers, and check the formulas against hand-worked cases: ABS's 6-week pump example, its Table 2, and the RCM worksheet's relief valve (15-year MTBF, a demand every 8 years, one multiple failure in 2,000 years tolerated: 43.8 days).

Embed this tool

Teaching this, or writing about it? Paste this code into your site, course page or intranet and the tool works right there. It is free, with no sign-up.

<iframe src="https://www.theleansuite.com/tools/maintenance-strategy-selector/embed" title="Maintenance strategy selector by LeanSuite" width="100%" height="1500" style="border:0;max-width:1000px" loading="lazy" allow="clipboard-write"></iframe>
<p style="font:13px/1.4 sans-serif"><a href="https://www.theleansuite.com/tools/maintenance-strategy-selector">Maintenance strategy selector</a> by LeanSuite</p>

FAQ

Maintenance strategy selector: common questions

More free lean tools

All tools

Pass it on

We want LeanSuite to be the best place on the internet for lean help. If this was useful, send it to someone on your team or in your network who needs it.