LeanInsight Library

Maintenance KPIs: What to Measure and Why

Lean Manufacturing Education

Lean Manufacturing Education

Master foundational lean principles and Toyota Production System concepts spanning 8 wastes, 5S, TPM, kaizen, value stream mapping, and standardized work.

Author

Aileen Nguyen

Aileen Nguyen

Content Architect

Vibhav Jaswal is a content architect who turns complex technical subjects into clear, well-organized knowledge systems. With a background in graphic design and project management, he focuses on breaking down intricate concepts and connecting them in ways that make sense to the reader, from first principles all the way through to practical application. His work spans educational content, visual resources, and product documentation. At LeanSuite, he applies this to lean manufacturing, building structured content that helps production teams understand and implement the tools and methods that drive operational improvement.

Articles by Aileen Nguyen

Published

Updated

Reading Time

16 mins

Maintenance KPIs are the measurable values that reveal whether a maintenance program is actually working, covering equipment reliability, planning discipline, cost, and workload in numbers that can be tracked over time and compared against a benchmark. Tracking the right five to ten of these consistently produces more actionable insight than tracking twenty superficially, since a smaller, well-chosen set of KPIs gets genuinely reviewed and acted on, while a sprawling dashboard tends to get glanced at without driving decisions.

The purpose of maintenance KPIs is not measurement for its own sake. Each one exists to answer a specific operational question: is equipment reliable, is maintenance proactive or reactive, is the maintenance team keeping pace with demand, and is the plant getting value from its assets. Selecting KPIs starts with knowing which of these questions matters most right now, not with copying a generic list.

This guide covers the core reliability, planning, workload, and value metrics worth tracking, how each is calculated, and how to select the right subset for a specific operation rather than trying to track everything at once.

Reliability Metrics: MTBF and MTTF

Reliability metrics answer the most fundamental maintenance question: how often does equipment actually fail.

The Two Core Reliability Measures

Mean Time Between Failures (MTBF). MTBF measures the average operating time between failures for repairable equipment, calculated as total operating time divided by number of failures. A machine running 1,000 hours with 5 failures has an MTBF of 200 hours. A higher MTBF indicates more reliable equipment, though the benchmark varies significantly by equipment type and should be tracked against the specific asset's own history rather than a generic industry number.

Mean Time To Failure (MTTF). MTTF applies specifically to non-repairable assets, components that get replaced rather than repaired when they fail, calculated as total operating time divided by number of failed units. Ten identical bulbs running a combined 10,000 hours before failing gives an MTTF of 1,000 hours per bulb. The distinction between MTBF and MTTF matters because applying the wrong metric to the wrong asset type produces a number that looks meaningful but answers the wrong question.

Key Insight: MTBF applies to repairable equipment and MTTF to non-repairable components, and confusing the two produces a technically calculated but practically meaningless number.

A Worked Calculation Example: MTBF and PMP for a Real Line

The formulas are easier to apply against real numbers than described abstractly. Consider a stamping line that ran for 720 hours last month, logging 4 unplanned failures, alongside 160 total maintenance hours for the same period, of which 130 were planned.

Calculating MTBF

MTBF is total operating time divided by number of failures: 720 hours divided by 4 failures equals an MTBF of 180 hours. This means the line runs an average of 180 hours between unplanned stoppages. Tracked month over month, a declining MTBF, say dropping to 150 hours the following month with the same failure count expected, signals a developing reliability problem worth investigating before it worsens further.

Calculating PMP

PMP is planned maintenance hours divided by total maintenance hours: 130 planned hours divided by 160 total hours equals a PMP of 81 percent. This sits just under the 85 percent benchmark considered strong, indicating the line is mostly proactive but has room to convert more of its remaining reactive hours into planned work. The 30 reactive hours, once broken down by cause, typically reveal exactly which equipment or failure mode is driving the gap.

Key Insight: Applying the formulas to real numbers, not just naming them, is what turns MTBF and PMP from abstract definitions into monthly figures a maintenance team can actually track and act on.

Planning and Reactivity Metrics

These metrics reveal whether a maintenance program is operating proactively or simply responding to failures as they occur.

Measuring the Planned-to-Reactive Balance

Planned Maintenance Percentage (PMP). PMP measures the share of total maintenance hours that were planned in advance rather than reactive, calculated as planned maintenance hours divided by total maintenance hours. A PMP of 85 percent or higher is generally considered strong, reflecting the proactive discipline covered in [Planned Maintenance: Optimization Strategies for Reliability].

Reactive Maintenance Percentage. The direct counterpart to PMP, calculated as reactive maintenance hours divided by total maintenance hours. World-class organizations typically target under 10 percent, since reactive work is both more expensive and more disruptive than the equivalent planned task.

Emergency Maintenance Percentage. A narrower, more urgent subset of reactive work, calculated as emergency maintenance hours divided by total maintenance hours. Under 5 percent is generally considered optimal, and a high or rising emergency percentage is one of the clearest early signals that a planned maintenance program has real gaps.

Key Insight: PMP, reactive percentage, and emergency percentage together reveal how proactive a maintenance program genuinely is, not just whether it exists on paper.

Workload and Compliance Metrics

These metrics track whether the maintenance team's capacity is keeping pace with the work identified, and whether scheduled work actually gets completed.

Tracking Capacity and Execution

Maintenance Backlog. Backlog measures pending or overdue maintenance work as a percentage of total available maintenance hours, or more practically, as a number of weeks of queued work. A backlog of one to two weeks is generally considered healthy, as covered in more depth in [Planned Maintenance: Optimization Strategies for Reliability].

Schedule Compliance. Schedule compliance measures the percentage of scheduled maintenance tasks actually completed on time, calculated as completed tasks divided by scheduled tasks. Compliance above 90 percent generally indicates disciplined planning and execution, while a lower number signals either overscheduling or capacity gaps.

Maintenance Technician Productivity. This metric tracks the share of technician hours actually spent on maintenance work versus total hours worked, calculated as completed maintenance hours divided by total hours worked. A productivity figure above 85 percent generally indicates efficient scheduling and minimal wasted technician time.

Key Insight: Backlog, schedule compliance, and technician productivity together reveal whether maintenance capacity is genuinely matched to the work the plant actually needs done.

Asset Value and Downtime Metrics

These metrics connect maintenance performance directly to the financial and production value of the equipment being maintained.

Connecting Maintenance to Asset Value

Equipment Downtime. Downtime measures total non-operational time as a percentage of total operating hours, calculated as downtime hours divided by operating hours. Under 5 percent is generally considered excellent, though the acceptable threshold varies by how critical the specific equipment is to production.

Remaining Asset Value (RAV). RAV evaluates an asset's current value against its original value after depreciation and maintenance costs, calculated as current value divided by original value. A declining RAV signals an asset approaching the point where replacement becomes more economical than continued maintenance investment.

Overall Equipment Effectiveness (OEE). OEE combines availability, performance, and quality into a single comprehensive score, and functions as the broadest measure of how effectively equipment is actually being used across the full library's OEE and equipment metrics coverage. World-class OEE is generally considered 85 percent and above.

Key Insight: RAV and downtime connect maintenance directly to financial decisions, while OEE remains the single broadest measure of equipment effectiveness overall.

Selecting the Right KPIs for a Specific Operation

Tracking every available maintenance KPI is rarely the right approach. The goal is a focused set, typically five to ten metrics, that genuinely gets reviewed and acted on.

Starting From the Objective, Not the Metric List

Selection should start with the specific operational objective, reducing downtime, improving reliability, or controlling maintenance cost, rather than starting from a list of available metrics and picking the most familiar ones. Equipment criticality also shapes selection: a plant with a small number of highly critical assets may prioritize MTBF and downtime closely, while an operation with many redundant, lower-criticality assets may weight planning and cost metrics more heavily.

Reviewing and Refining the Set Over Time

The selected KPI set should be revisited periodically rather than treated as permanent, since operational priorities shift as a maintenance program matures. A plant early in its planned maintenance journey may prioritize PMP and reactive percentage to establish proactive discipline first, then shift emphasis toward MTBF and RAV once that discipline is established.

Key Insight: The right KPI set starts from the operational objective and equipment criticality, not from copying a generic list, and should be revisited as the maintenance program matures.

A Worked Example: Two Plants, Two KPI Sets

The selection principle is easier to see applied to two contrasting plants rather than described abstractly: an early-stage manufacturer just formalizing maintenance, and a mature plant with an established planned maintenance program already in place.

The Early-Stage Plant

A plant still transitioning away from purely reactive maintenance gets the most value from tracking planned maintenance percentage, reactive maintenance percentage, and maintenance backlog first. These three metrics directly measure whether the transition is actually happening, since the immediate goal at this stage is building proactive discipline, not yet optimizing reliability at a granular level. MTBF and RAV would be premature priorities here, since there is limited planned maintenance history yet to make those numbers meaningful.

The Mature Plant

A plant with planned maintenance percentage already above 85 percent and reactive work under 10 percent has largely solved the proactivity question, so its KPI priorities shift toward MTBF, schedule compliance, and RAV. These metrics answer a more advanced set of questions: is the existing program actually improving reliability, is capacity keeping pace with demand, and which assets are approaching a genuine replacement decision.

Key Insight: The same KPI list serves different plants differently, with an early-stage program prioritizing proactivity metrics and a mature program shifting toward reliability and asset value metrics.

Within the Lean System

Connection to Lean Principles

Maintenance KPIs operationalize the lean principle of managing by fact rather than assumption, giving maintenance decisions a measurable basis rather than relying on impression or habit. This connects directly to [Total Productive Maintenance: A Complete Manufacturing Guide]'s broader emphasis on outcome-based measurement over activity-based measurement.

Connection to Lean Tools

These KPIs measure the results of the strategies covered throughout this cluster: PMP and reactive percentage measure [Planned Maintenance: Optimization Strategies for Reliability] in action, while MTBF and downtime measure the effectiveness of the strategy assignments made through [Reliability-Centered Maintenance: A Practical Guide].

Connection to Continuous Improvement

Tracking these KPIs over time is what feeds the [PDCA Cycle: The Foundation of Continuous Improvement] at the maintenance program level, turning a static set of numbers into an ongoing check phase that reveals whether recent changes to maintenance strategy are actually producing the intended reliability and cost improvements.

Frequently Asked Questions

Q: What are the most important maintenance KPIs to track?

MTBF, planned maintenance percentage, maintenance backlog, and equipment downtime cover reliability, proactivity, workload, and production impact respectively, making them a strong starting set before adding more specialized metrics based on specific operational priorities and the criticality of the equipment actually involved.

Q: How many maintenance KPIs should a plant track?

Typically five to ten metrics total. Tracking more than that tends to dilute attention and produces a dashboard that gets glanced at rather than genuinely reviewed and acted on, which directly undermines the actionable insight KPIs are supposed to provide in the first place.

Q: What is the difference between MTBF and MTTF?

MTBF applies to repairable equipment and measures average time between failures. MTTF applies specifically to non-repairable components that get replaced rather than fixed when they fail. Applying the wrong metric to the wrong asset type produces a technically calculated but practically meaningless figure overall.

Q: What is a good planned maintenance percentage?

A PMP of 85 percent or higher is generally considered strong, reflecting a maintenance program where most work is genuinely scheduled in advance rather than reactive. Correspondingly, world-class organizations typically target their reactive maintenance percentage at under 10 percent overall.

Q: How often should maintenance KPI selection be reviewed?

Periodically, rather than treated as a fixed and permanent list, since operational priorities genuinely shift as a maintenance program matures over time. An operation building proactive discipline may prioritize PMP and reactive percentage initially, then shift emphasis toward reliability and asset value metrics later.

Harnessing technology for maintenance KPI optimization

Leveraging tools and software to create, track, and analyze maintenance KPIs is crucial for maintaining operational excellence and efficiency.

Advanced solutions like LeanSuite's KPI Builder provide a comprehensive platform that empowers you to design custom KPIs tailored to your specific maintenance needs. This tool not only facilitates the creation and management of diverse KPIs, but also offers robust tracking and analytical capabilities.

By utilizing such technology, you can gain real-time insights into your maintenance operations, identify trends, and make data-driven decisions to optimize performance. The ability to monitor and evaluate KPIs effectively ensures that maintenance strategies align with broader business goals. Ultimately, enhancing asset reliability, reducing downtime, and driving continuous improvement.

LeanSuite: A complete lean manufacturing software

Schedule Demo
Blog Banner