Production downtime falls when a factory defines every event consistently, separates failure frequency from recovery time and prioritizes losses by their effect on throughput. Buying a CMMS or adding sensors can support the work, but neither tool removes a cause by itself.
A useful event timeline starts when production capability is lost and ends with the first good unit after a stable restart. Between those points, the plant should distinguish detection, reporting, waiting, diagnosis, parts or tools, repair, test, restart and quality recovery. Each interval requires a different improvement.
What counts as production downtime
Downtime should be measured against a defined production requirement and a consistent boundary. For many discrete processes, a practical boundary is the loss of required capability to the first good unit produced after stable restart.
The plant must decide what the clock measures. Calendar time, scheduled production time and operating time answer different questions. A machine stopped on a weekend with no demand may not create lost production. The same stop during a constrained shift can have a high cost.
Separate planned and unplanned events. Planned maintenance, cleaning and changeover may still consume capacity even though they are scheduled. Unplanned failure, missing material, quality hold, software fault and absent operator may all stop output but require different owners.
Microstops deserve a rule. A threshold such as five minutes can make reporting easier, but stops below the threshold do not disappear. They may be captured automatically as minor stops or speed loss. If operators repeatedly reset a fault for 20 seconds, the accumulated loss can exceed one dramatic breakdown.
The end point matters. Ending downtime when maintenance hands the machine back can hide warm-up, adjustment, inspection and scrap. Ending at the first good unit creates a cross-functional measure of recovery. For continuous processes, use an equivalent stable-rate and quality condition.
Document the definition beside every dashboard. Comparisons are misleading when one line includes changeovers and another excludes them, or one plant stops the clock at repair completion while another waits for first good product.
Which downtime metrics should a factory use
MTBF measures how often repairable failures occur. MTTR measures how long restoration takes. Inherent availability combines the two, but it is not the same as overall equipment effectiveness or saleable output.
Mean time between failures is operating time divided by the number of relevant failures. The failure definition and population must be stable. Do not combine unrelated assets or count minor stops in one month and exclude them in the next.
Mean time to repair is often used loosely. It can mean hands-on repair, restoration from stop to function, recovery to normal production or resolution of an incident. The factory should state the start and end events. For throughput decisions, restoration to first good unit is often more useful than wrench time.
Availability can be estimated as MTBF divided by MTBF plus MTTR when assumptions fit. This relationship explains two distinct routes: increase the time between failures or reduce the time to restore. It does not capture speed loss, scrap, planned stops or demand.
| Metric | Basic definition | Decision supported | Common misuse |
|---|---|---|---|
| Downtime hours | Time unavailable within defined schedule | Size and location of loss | Mixed start and end rules |
| Failure count | Number of defined functional failures | Reliability and recurrence | Counting symptoms as separate failures |
| MTBF | Operating time divided by failures | Failure-frequency improvement | Comparing unlike assets or periods |
| MTTR | Restoration time divided by repairs | Recovery improvement | Calling only hands-on time MTTR without disclosure |
| Availability | MTBF divided by MTBF plus MTTR | Reliability-maintainability balance | Presenting it as OEE |
| Good output loss | Expected good units minus actual good units attributable to stop | Throughput and financial impact | Using theoretical maximum without demand context |
Use distributions as well as averages. A median repair of 20 minutes with a few eight-hour events needs a different response from consistently long repairs. Track the 90th percentile, maximum and recurrence by failure mode.
How should downtime data be collected and classified
Downtime data should distinguish the symptom, failure mode, cause, waiting state, corrective action and restart result. A single code called “mechanical” cannot support root-cause work.
Automatic machine signals provide precise timestamps but limited meaning. Operator input adds context but can be inconsistent. Combine them: use the control system for start and state, then require a short structured classification by the responsible role.
Create a code hierarchy with few top-level categories and controlled detail. Examples include equipment failure, material, quality, changeover, staffing, utilities and external dependency. Equipment failure can then branch into asset, failure mode and cause. Avoid hundreds of codes that users cannot select reliably.
The event record should include:
- asset and line identity;
- start, detection, report, work start, repair complete and first-good timestamps;
- product, order, shift and operating state;
- symptom and functional failure;
- confirmed failure mode and cause when known;
- parts, tools and skills used;
- action taken and whether it was temporary or permanent;
- restart scrap and time to stable production;
- person confirming closure.
Unknown is a valid temporary classification. Forcing an immediate cause encourages guesses. A high-impact event can enter root-cause analysis after evidence is collected. The final code should be updated without erasing the original observation.
Audit data quality weekly during implementation. Look for missing timestamps, excessive “other,” identical durations, reopened events and mismatches between work orders and machine states. A dashboard built on poor codes creates false precision.
How should downtime losses be prioritized
Downtime priorities should combine duration, frequency, financial effect and impact on the system constraint. One Pareto chart is not enough.
A duration Pareto identifies the failure modes that consume the most hours. A frequency Pareto shows chronic events. A cost view adds repair, scrap and lost contribution. A constraint view asks whether the event actually limited saleable output.
Suppose fault A occurs twice and consumes ten hours. Fault B occurs 120 times and consumes six hours. Fault A leads the duration chart. Fault B may still deserve priority if every reset interrupts the bottleneck, distracts operators and produces startup scrap. Conversely, a long stop on a machine with spare capacity may have little effect on customer output.
Estimate throughput impact using the best alternative production path. If another line can absorb the volume, include transfer, overtime and efficiency costs rather than full lost sales. If the constraint has no recovery capacity, value an hour using finance-approved contribution, not revenue.
After ranking, select a small number of problems. One high-impact intermittent failure may need formal root-cause analysis. A chronic microstop may need direct observation and a simple engineering change. A long waiting interval may be solved by spare-parts policy or escalation rather than machine redesign.
How can a factory improve MTBF
MTBF improves when the plant removes failure causes and conditions that accelerate degradation. More inspections alone do not necessarily increase reliability.
Start with repeat failures. Confirm the failure mode through parts, photographs, measurements and control history. Use a structured method such as fault-tree analysis, five whys or cause-and-effect analysis according to complexity. Verify the proposed cause before implementing an expensive solution.
Common MTBF actions include correcting contamination, alignment, lubrication, cooling, electrical quality, software logic, overload, poor installation and weak incoming components. Standardize operating conditions and detect deviations. Review whether preventive tasks are technically justified and performed without introducing defects.
Design changes may be necessary. Improve access, sealing, supports, component rating or control limits. Manage the change and verify that it does not transfer risk elsewhere. A stronger component can expose another weak point.
Use preventive versus predictive maintenance to assign the right strategy to each asset. Condition monitoring supports MTBF only when it enables action before functional failure or reveals a condition that can be removed.
How can a factory improve MTTR
MTTR improves when detection, escalation, diagnosis, access, parts, repair and restart are faster and more reliable. The largest opportunity may be waiting time rather than technician speed.
| Loss interval | Typical improvement lever | Owner evidence |
|---|---|---|
| Detection and reporting | Reliable state signal and escalation rule | Alarm history and response timestamp |
| Waiting for access or permit | Preplanned safe access and clear authorization | Permit and queue data |
| Diagnosis | Fault history, test points and decision guide | Diagnostic time by failure mode |
| Waiting for parts or tools | Critical-spares policy and prepared kit | Stock availability and retrieval time |
| Hands-on repair | Access, modular replacement and standard work | Repair-time distribution |
| Test and restart | Defined checks and prepared materials | First-good timestamp and startup scrap |
Break the event into intervals. If a machine waits 40 minutes for notification and ten minutes for repair, remote alerting or escalation may matter more than a faster tool. If diagnosis dominates, provide fault history, schematics, test points and decision trees. If parts dominate, review critical-spares location and replenishment.
Standard work can prepare tools, permits and replacement kits. Improve physical access during design reviews. Use connectors and modular replacement where appropriate. Train more than one person for critical recovery tasks and establish supplier escalation for problems beyond local capability.
Do not shorten MTTR by bypassing safety. The OSHA control of hazardous energy resource explains requirements around hazardous energy control. Local law, risk assessment and authorized procedures govern the work. A restart that exposes people or creates a quality escape is not an improvement.
Define restart standards. Check guards, parameters, material state and quality. Measure time to first good unit and capture startup scrap. A repaired machine that needs an hour of adjustment has not completed recovery when the motor first turns.
What role do TPM predictive maintenance and CMMS play
TPM, predictive maintenance and CMMS are useful when they address a defined loss and connect data to action. They are not substitute names for a reliability program.
Total productive maintenance involves operators and maintenance in basic condition, inspection, problem elimination and equipment effectiveness. It can reveal contamination, loose components and abnormal conditions before failure. The boundaries of operator work must remain safe and authorized.
Predictive maintenance can improve timing for detectable failure modes. It is not suited to every asset and depends on response. A good alert ignored for three days does not reduce downtime.
CMMS supports asset history, planning, work orders, parts and feedback. Its value depends on code quality and use. Adding mandatory fields can improve data or encourage meaningless entries. Configure the minimum information needed for decisions and audit it.
What does a documented TPM case show
A single case can show the mechanisms of improvement, but its percentages should not become a universal benchmark.
A NIST MEP case involving Leggett and Platt Aerospace reported about a 22 percent productivity increase, an OEE change from 39 to 45 percent and USD 250,000 in avoided investment after TPM work on a problematic CNC line.
“The biggest standout of this process was implementing a practice of TPM and applying it to our other machines.”
Alec Martone, Manufacturing Engineer at Leggett and Platt Aerospace, made this statement in the case. The transferable lesson is the documented practice and expansion to other equipment, not an assumption that every TPM project creates the same percentage.
A 90-day production downtime reduction plan
Days 1 to 14: agree definitions, clocks and owners. Select one line. Validate automatic states against observation. Establish baseline hours, events, MTBF, restoration intervals, first-good recovery and lost output.
Days 15 to 30: clean classifications and build four views: duration, frequency, financial effect and constraint effect. Select two losses with different mechanisms. Confirm the problem at the workplace.
Days 31 to 60: run focused countermeasures. One stream should increase time between failures. The other should reduce waiting or recovery time. Define expected change, owner and verification period before action.
Days 61 to 90: confirm results across normal product and shift variation. Update standard work, preventive tasks, spares and training. Remove temporary controls that are no longer needed. Review whether the improvement moved the constraint or created another loss.
At day 90, decide whether to standardize, extend or investigate further. Do not scale a dashboard before the event definition and response process are stable.
Frequently asked questions
What is a good MTBF
There is no universal good MTBF. It depends on asset function, duty, consequence and design. Compare the same failure definition over time and with engineering or supplier evidence.
Does planned changeover count as downtime
It can count as planned downtime or setup loss when it consumes required capacity. Keep it separate from unplanned failure so the correct method and owner remain visible.
How should microstops be recorded
Capture them automatically when possible and aggregate by fault or state. Use a threshold that supports response, but do not discard shorter events from loss calculations.
Can a CMMS reduce downtime by itself
No. A CMMS can improve planning, records and response. Downtime changes only when people use the information to remove causes or restore production more effectively.
Is availability the same as OEE
No. Availability is one OEE component. OEE also includes performance and quality. For capacity decisions, good output and the system constraint may be more useful than a local OEE number.
Sources
- Maintenance Costs and Advanced Maintenance Techniques in Manufacturing Machinery, NIST.
- Monitoring Diagnostics and Prognostics for Manufacturing Operations, NIST.
- Total Productive Maintenance, Lean Enterprise Institute.
- What Is Problem Solving, ASQ.
- MTTR vs MTBF, IBM.
- Control of Hazardous Energy, OSHA.
- TPM Reduces Equipment Downtime and Lost Capacity, NIST MEP, 2022.




