Free tool
Maintenance strategy selector
To choose how to maintain a failure mode, ask three questions about its consequences: would anyone notice it, could it hurt someone or breach an environmental limit, and does it stop output. Then try the tasks in order: a condition check if there is a warning sign that lasts longer than the time you need to act, a time-based replacement only if failures rise at a known age, and otherwise the default for the consequence: redesign where someone could be hurt, run to failure where it only costs money, a failure-finding test where the failure is hidden. That is the order of the RCM decision diagram Nowlan and Heap published in 1978. A bearing with 6 weeks of vibration warning is checked every 3 weeks; a relief valve with a 20-year MTBF that must be 99.5% available is tested every 73 days.
Check intervalequalsthe smaller of P-F interval divided by 2 and (P-F interval − time to act)
Failure-finding intervalequals2 × (1 − A) × MTBF
Failure-finding intervalequals2 × MTBF × MTED divided by MMF
One failure mode at a time: a specific way the asset fails, such as "drive-end bearing wears", not "the compressor fails". The questions follow the order of the RCM decision diagram. "Not sure" takes the cautious default and marks the answer provisional.
Recommended strategy
Condition-based
Check for the warning sign at a fixed interval and plan the repair when the sign is found.
Consequence
Operational
The failure stops or slows output, or harms quality or delivery.
Interval basis
Check every 3 weeks
Half the P-F interval of 6 weeks; leaves 3 weeks to act, 1 week needed.
Why: the path through the questions
- 1Evident in normal work? YesThe failure is evident, so its own consequences decide what is worth doing.
- 2Safety or environment? NoNo safety or environmental consequence, so the choice is an economic one.
- 3Affects output or quality? YesOperational: a task is worth doing if it costs less than the repairs plus the lost output it prevents.
- 4Warning sign before failure? YesA condition check may work, if the warning lasts long enough.
- 5P-F interval and time to act 6 weeks warning, 1 week to actCheck every 3 weeks (half the P-F interval), which leaves at least 3 weeks to act.
- 6Warning time consistent? YesThe warning time can be relied on.
- 7Check practical? YesThe check can be done at that interval.
- 8Checks cheaper than failures? YesThe checks pay for themselves.
Data to collect next
- Record each reading against its alert limit. Shorten the interval if the sign is found late; lengthen it if readings stay flat over several checks.
Supports a decision for one failure mode; it does not replace an RCM analysis with the people who run and maintain the asset. For the full study, use the RCM worksheet.
Summary table
Add each failure mode as you finish it, then copy the table into your PM plan or CMMS import sheet. The example rows are the worked example below.
| Asset | Failure mode | Strategy | Interval basis | Owner | Remove |
|---|---|---|---|---|---|
| Air compressor C-2 | Motor drive-end bearing wears | Condition-based maintenance | Check every 3 weeks | Reliability technician | |
| Air compressor C-2 | Pressure relief valve sticks shut | Failure-finding test | Test every 0.2 years (73 days) | Maintenance planner | |
| Air compressor C-2 | Panel indicator lamp burns out | Run to failure | No scheduled task; repair on failure | Operator |
Next step
Want to schedule and track these tasks?
Bring this result to a 45-minute demo and we will show where it fits in LeanSuite's Professional Maintenance Tags.
How to use it
One failure mode at a time
1. Choose the asset, then its failure modes. An equipment criticality analysis picks the asset worth the time. A failure mode is a specific cause, such as "drive-end bearing wears" or "relief valve sticks shut", taken from work orders and from the technicians who fix it.
2. Answer the consequence questions. The first one, whether anyone would notice the failure in normal work, is the one most often answered wrongly: protective devices and standby units can sit failed for months without a sign.
3. Answer the task questions from evidence. Warning signs and warning times come from the technicians and from condition readings, failure ages from your work orders. "Not sure" takes Nowlan and Heap's cautious default answers: hidden, safety, operational, a condition check assumed workable, a scheduled replacement assumed not until there is data. The result is then marked provisional.
4. Read the reasoning, then set the interval. Round the interval down to a slot in your schedule you can keep, and shorten it for a failure that could hurt someone.
5. Add it to the table and do the next one. Copy the table into your PM plan. Review each interval as findings come in: a sign found late means a shorter interval; readings that stay flat over several checks allow a longer one.
Worked example
Illustrative numbers, not a benchmarkThree failure modes of a rotary screw air compressor that supplies a line. The warning time, the MTBF and the availability target are illustrative, not data from a real machine.
- 1Drive-end bearing wears. Evident (it trips the motor), no safety consequence here, and the line loses air: operational. Vibration analysis picks up the defect about 6 weeks before failure, and getting the bearing and a stop takes 1 week.
- 2Check every 6 ÷ 2 = 3 weeks. The worst case, a defect that appears just after a check, leaves 6 − 3 = 3 weeks to act, more than the 1 week needed. The checks cost less than a failed bearing and a stopped line: condition-based maintenance every 3 weeks.
- 3Pressure relief valve sticks shut. Nobody notices until the pressure control also fails, and that multiple failure could hurt someone: hidden, safety. No warning sign without a test and no known wear-out age, so the default is a failure-finding test, done on a test bench.
- 4With a valve MTBF of 20 years and 99.5% availability set by the site: 2 × (1 − 0.995) × 20 = 0.2 years = 73 days. 73 days is 1% of the MTBF, well under the one-tenth limit: test every 73 days, scheduled every 10 weeks.
- 5Panel indicator lamp burns out. Evident, harmless and costs only a lamp: non-operational. No warning sign, failures at random: run to failure, with spare lamps in stock.
Why the questions come in this order
The tool follows the RCM decision diagram in Nowlan and Heap's report for the US Department of Defense (Exhibit 4.4). Consequences come first, because they decide what "worth doing" means:
- Safety or environment: a task has to bring the risk to a level the site accepts, whatever it costs. If none does, redesign is required.
- Operational: a task has to cost less than the repairs plus the lost output it prevents.
- Non-operational: a task has to cost less than the repairs alone.
- Hidden: a task has to give the protective function the availability you need. The default is a failure-finding test.
Within each branch a condition check is tried first, because it replaces a part only when the part needs it; Nowlan and Heap call it the most desirable task whenever it applies. Environmental and legal consequences sit with safety, as in later RCM standards such as MIL-STD-3034A and the NASA RCM guide.
The six failure patterns
United Airlines plotted the chance of failure against age for its aircraft components and found six shapes. Nowlan and Heap published them in 1978 (Exhibit 2.13), with the share of the items studied in each:
- A. Bathtub, 4%High when new, then low and steady, then a wear-out zone.
- B. Wear-out, 2%Steady or slowly rising, then a pronounced wear-out zone.
- C. Slowly rising, 5%Rises gradually with age, with no clear wear-out age.
- D. Low, then steady, 7%Low when new or just repaired, then rises quickly to a steady level.
- E. Random, 14%The same chance of failure at every age.
- F. Early life, 68%High when new or just repaired, then steady or very slowly rising.
In the report's words, some 89 percent of the items had no wear-out zone, so an age limit could not improve them; 11 percent (A, B and C) might benefit from one. The NASA RCM guide (2008) adds a Swedish study (1973) and a US Navy study (1982) and puts random failures at 77 to 92 percent across the three.
These are aircraft and naval items, not plant equipment, and we found no comparable published study for plants that we could check. What carries over is the test: replacing on a schedule helps only a failure mode with a clear wear-out age that most units reach. Early-life failures (pattern F) point to installation, start-up and repair quality, and the NASA guide notes that a scheduled overhaul often adds more of them.
Axes: chance of failure (vertical) against age since new or since repair. Some later texts letter the patterns in a different order; these are Nowlan and Heap's.
The P-F interval and the check interval
Nowlan and Heap define a potential failure as "an identifiable physical condition which indicates a functional failure is imminent". A condition check works only if that condition can be detected, the time from it to the failure is reasonably consistent, and there is time to act.
- Check at no more than half the P-F interval (ABS, 4.1). With a 6-week warning, check every 3 weeks.
- Leave time to act. In the worst case the sign appears just after a check, so the time left is the P-F interval minus the check interval. If that is shorter than the time you need to plan and do the repair, shorten the interval. If the warning is not longer than the time to act, no check works.
- Shorten it further for higher-risk failure modes or when the P-F estimate is a guess (ABS, 4.1). Once a reading passes its alert limit, the NASA guide says to cut the monitoring interval to between a third and a quarter of what it was.
- Learn the real P-F interval. ABS gives the example of pumps whose weekly vibration readings caught defects that then ran 6 to 8 weeks before repair without failing: the P-F interval is at least 6 weeks, so the check can move to every 3 weeks.
The failure-finding interval
A protective device that fails at random and is tested every T is failed, on average, for about T ÷ (2 × MTBF) of the time. Set that equal to the unavailability you can accept and solve for T: T = 2 × (1 − A) × MTBF. If you know how often the device is called on (MTED) and how rarely you can accept the multiple failure (MMF), the unavailability is MTED ÷ MMF.
| Unavailability accepted | Test interval |
|---|---|
| 0.0001 (99.99% available) | 0.02% of MTBF |
| 0.001 (99.9%) | 0.2% of MTBF |
| 0.01 (99%) | 2% of MTBF |
| 0.05 (95%) | 10% of MTBF |
The unavailability you accept is your site's risk rule, not a standard. ABS lists the assumptions: random failures, failure rate × interval under 0.1, test and repair times short, and the multiple failure possible only from that one demand. The test itself must not create the hazard, and should prove the whole protective function, not one part of it.
What this tool does not do
- It does not define functions or list failure modes. An RCM analysis does that with the people who run and maintain the asset; the RCM worksheet holds the whole study.
- It does not decide what is safe. Whether a failure could hurt someone, and what risk the site accepts, are decisions for the site and its safety professionals.
- It does not override a law, code or insurer. Where one sets an inspection or test interval, as is common for pressure relief devices, that interval applies whatever this tool says.
- It does not supply failure data. Every warning time, MTBF and cost is yours; the tool only shows what follows from them.
- No claim is made that it meets SAE JA1011, the standard that sets what a process must include to be called RCM.
Sources and how this tool was checked
- F. Stanley Nowlan and Howard F. Heap, Reliability-Centered Maintenance (US Department of Defense, 1978, report AD-A066579): potential failure (section 2.1), the six patterns (section 2.8, Exhibit 2.13), the task criteria (chapter 3), the decision diagram and default answers (Exhibits 4.4 and 4.5).
- American Bureau of Shipping, Guidance Notes on Reliability-Centered Maintenance (2004, updated 2018): section 4 for the P-F and check intervals, section 5 for the failure-finding interval and Table 2.
- NASA, Reliability-Centered Maintenance Guide for Facilities and Collateral Equipment (2008): chapter 3 for the four outcomes, section 4.1.5 for the three failure-pattern studies, page 4-3 for shortening the interval after an alert.
- John Moubray, Reliability-centred Maintenance (RCM II), 2nd ed. (1997), the usual reference for the P-F interval and both interval rules, is not free to read and was not used to check this page.
Unit tests walk every combination of answers to make sure each ends in a recommendation, check the paths above against the decision diagram and the default answers, and check the formulas against hand-worked cases: ABS's 6-week pump example, its Table 2, and the RCM worksheet's relief valve (15-year MTBF, a demand every 8 years, one multiple failure in 2,000 years tolerated: 43.8 days).
Embed this tool
Teaching this, or writing about it? Paste this code into your site, course page or intranet and the tool works right there. It is free, with no sign-up.
<iframe src="https://www.theleansuite.com/tools/maintenance-strategy-selector/embed" title="Maintenance strategy selector by LeanSuite" width="100%" height="1500" style="border:0;max-width:1000px" loading="lazy" allow="clipboard-write"></iframe>
<p style="font:13px/1.4 sans-serif"><a href="https://www.theleansuite.com/tools/maintenance-strategy-selector">Maintenance strategy selector</a> by LeanSuite</p>How LeanSuite helps
In LeanSuite, professional maintenance tags open a work order when an operator tags an issue or a scheduled PM is due, build a searchable repair history for every asset, and show MTTR and MTBF by machine or line. That record is the evidence for the failure ages and MTBF this tool asks about.
See Professional Maintenance TagsRead more
- Free template: RCM worksheet (failure mode and task selection)
- Free template: Equipment criticality analysis
- Free template: Preventive maintenance checklist template
- Free template: Critical spare parts list and stocking levels
- Reliability-centered maintenance (RCM) in the lean glossary
- Reliability-Centered Maintenance: A Practical Guide
- P-F interval in the lean glossary
- Condition-based maintenance in the lean glossary
- Failure-finding task in the lean glossary
- What are Five Types of Maintenance Strategies?
- Preventive vs Predictive Maintenance: Key Differences
FAQ
Maintenance strategy selector: common questions
More free lean tools
All tools- OEE calculatorOverall equipment effectiveness from shift time, downtime, ideal cycle time and part counts.
- Downtime cost calculatorWhat an hour, an event and a year of unplanned downtime cost, from lost output, idle labour, scrap and overtime.
- SMED changeover calculatorTime saved by moving changeover steps from internal to external, capacity freed per week and the smaller batch it allows.
- Takt time calculatorThe pace a line must hit to meet customer demand, from shift time, breaks and daily demand.
- Cycle time calculatorCycle time against takt, units per hour, capacity per shift and the gap to daily demand.
- Line balancing calculatorLine efficiency, balance delay and the minimum number of stations from station times and takt.
- Bottleneck calculatorThe step that limits your line, from each step's cycle time, machines and uptime, with daily capacity and the gap to demand.
- First pass yield (FPY) calculatorFirst pass yield at each step and rolled throughput yield (RTY) across the whole process.
- Cost of poor quality (COPQ) calculatorTotal COPQ and cost of quality, and each as a share of revenue, from the four quality cost categories.
- Process capability (Cp, Cpk) calculatorCp and Cpk against your specification limits, plus Pp and Ppk from pasted measurements, with a distribution chart.
- DPMO and sigma level calculatorDPMO, defects per unit, yield and sigma level from the defects found, units checked and opportunities per unit.
- MTBF and MTTR calculatorMean time between failures, mean time to repair and availability from operating time, failures and repair time.
- Lead time calculatorManufacturing lead time from WIP and throughput (Little's Law), and the share of it that adds value.
- Throughput time calculatorThroughput time from process, inspection, move and wait time, the share of it that adds value, and units per hour.
- Kaizen savings calculatorYearly savings, payback in months and first-year ROI of an improvement, from time saved, scrap avoided and its costs.
- TEEP calculatorTotal effective equipment performance: OEE multiplied by how much of all calendar time you plan to run.
- Overall labor effectiveness (OLE) calculatorHow well paid labor hours turn into good output: availability, performance and quality of the crew.
- Inventory turnover and days on hand calculatorHow many times stock turns over in a period and how many days of it you hold, from COGS and average inventory.
- Safety stock calculatorSafety stock and reorder point from the service level you want, average demand and lead time, and how much each varies.
- Scrap rate calculatorScrap rate, the cost of scrap per period and per year, and what reaching a target rate would save.
- TRIR and DART rate calculatorOSHA recordable and DART incident rates from your 300A case counts and hours, and what one more case does to them.
- NIOSH lifting equation calculatorRecommended weight limit and lifting index for a two-handed lift, with every multiplier shown so you can see what to redesign.
- Time study sample size calculatorHow many cycles of an element to time, from your first readings, with the t formula, the ILO formula and the range method.
- Noise dose and TWA calculatorDose and 8-hour TWA for a shift of noise levels under OSHA and NIOSH, which limits are reached and how much loud time to cut.
- Repair or replace calculatorEquivalent annual cost of keeping and repairing an old machine against buying a new one, with each one's economic life and the break-even figures.
- Compressed air leak cost calculatorAir lost, kW and cost a year for each tagged leak from its size and the pressure, ranked so the biggest gets fixed first.
- Machine hour rate calculatorWhat one productive hour of a machine costs, with and without the operator, the cost per part and how utilisation moves the rate.
- Lean maturity self-assessmentScore 24 statements across eight dimensions and get a result band for each, an overall band and where to start.
- Manufacturing ROI calculatorAnswer a few questions about your shop floor, shifts and downtime to estimate your potential ROI with LeanSuite.

