QE Skill Enhancement (27) | Reliability Basics: Failure Rate, MTBF, and the Bathtub Curve

By: QTank Published: 10/7/2026 Views: 11
Current rating: ★★★☆☆ Rate this Equivalent to 8 ratings

1. How "MTBF 50,000 Hours" Was Misinterpreted as "Can Be Used for Over 5 Years"

A manufacturing company participated in a bid for an industrial control cabinet project. The customer's technical requirements stated "MTBF of the entire unit ≥ 50,000 hours." The R&D engineer immediately converted this: 50,000 hours ÷ 24 ÷ 365 ≈ 5.7 years, and replied "met." Three months later, during a technical clarification meeting, the quality engineer discovered three issues:

  1. The customer's reference was the average time to failure under actual operating conditions, not the bench data from a laboratory under normal temperature and no load.
  2. The value cited in the bid document came from the manufacturer's manual for a certain key component, while the entire unit consists of over a dozen modules in series. The failure rate adds up, and the MTBF is significantly reduced.
  3. Both parties had completely different understandings of the relationship between "MTBF and lifespan" — the customer assumed that "less than 5.7 years of use is unacceptable," while the company interpreted it based on an exponential distribution, believing that "average" means a significant number of units would fail much earlier than 5.7 years. If the quality engineer (QE) had not intervened, the bid commitment would have become a hidden risk that could not be fulfilled in production.

2. Statistical Logic of Three Quantities: Failure Rate, Reliability, and MTBF

The premise of reliability discussions is that "failures are random." During the random failure period, the failure rate λ is approximately constant, and the lifespan follows an exponential distribution. The relationship between reliability and time is:

R(t) = e^(−λt)

Three essential concepts to establish:

  1. λ is not a percentage but a failure intensity per unit time. Common units are FIT, where 1 FIT = 10⁻⁹/hour, meaning 1000 FIT = 1 failure per million hours.
  2. MTBF = 1/λ. This is a conversion relationship, not a lifespan definition; MTTF is used for non-repairable products, while MTBF is used for repairable products. For repairable products, MTBF = cumulative operating time ÷ number of failures, where the denominator is the number of failures rather than the number of units.
  3. R(MTBF) = e⁻¹ ≈ 36.8%. This is the most easily overlooked point — when the cumulative operating time equals MTBF, theoretically, only about 37% of the units are still working without failure.

The bathtub curve describes the three stages of failure rate over cumulative time: early failure period (decreasing failure rate due to manufacturing defects, material mixing, and assembly damage), random failure period (constant failure rate due to random stress), and wear-out failure period (increasing failure rate due to wear, aging, and fatigue). The strategies for these three stages are not interchangeable, so judgment is more important than calculation.

3. Practical Five Steps: From Indicator Definition to Acceptance Criteria

Step One: Lock Down the Indicator Definition. It must be clearly defined in the technical agreement with four elements: operating conditions (temperature/humidity/vibration/duty cycle), operating state (continuous or intermittent, including storage period or not), statistical scope (including repair time and planned downtime or not), and failure definition (only functional loss or including performance drift beyond limits). Criterion: When seeing "MTBF," one must immediately answer "under what conditions, what is being counted, and who is the denominator." If unable to answer, the definition is not locked down.

Step Two: Recognize Data Sources and Sample Size Criteria. Data from the manufacturer's manual can only be used for preliminary estimation, not as acceptance evidence. When estimating using bench or field data:

  • If failures occur, use the point estimate MTBF = ΣT ÷ r (ΣT is the cumulative operating hours, r is the number of failures).
  • If there are zero failures (r = 0) during the test period, the one-sided confidence lower limit is 2T ÷ χ²(2r+2, 1−C). Converted into an engineering criterion: at a 90% confidence level, it is necessary to accumulate approximately 2.3 times the target MTBF without any failures to claim "reaching the target"; at a 60% confidence level, about 0.92 times. In other words, "running for a week without breaking" statistically proves almost nothing.

Step Three: Use Data to Determine the Current Stage of the Bathtub Curve. Do not rely on intuition to label stages. Criterion: Use the Weibull shape parameter β to fit — β < 0.8 indicates the early failure stage, 0.8~1.2 indicates the random failure stage, and β > 1.2 indicates the wear-out stage. If the sample size is insufficient for fitting, use the trend test of cumulative failures and cumulative time (Laplace test) to assist in judgment. Incorrect stage determination will lead to all subsequent actions being wrong.

Step Four: Select Strategies Based on the Stage.

  • Early Failure Stage: The action is screening, not adding redundancy. Criterion: ESS (environmental stress screening) coverage rate (temperature cycling + vibration + aging) should be 100%, and the failure rate should decrease by at least one order of magnitude after screening (e.g., from 5 failures per thousand hours to less than 0.5 failures per thousand hours).
  • Random Failure Stage: The action is derating and redundancy. Criterion: Component derating (voltage ≤ 80% rated, power ≤ 70% rated, junction temperature with at least 25°C margin); after adding redundancy, adjust the system failure rate using a parallel model and separately evaluate common cause failures (same batch, same power supply, same vibration source).
  • Wear-Out Failure Stage: The action is preventive replacement and management of life-limited components. Criterion: Set the replacement cycle according to the B10 life (time when 10% of the units fail), generally taking 50%~70% of B10 as the replacement point; key life-limited components (bearings, seals, electrolytic capacitors, belts) should be included in the control plan and traced by batch.

Step Five: System-Level Conversion and Reporting. For a series system, reliability R_sys = Π R_i, and the system failure rate is approximately nλ when n identical units are in series — this is why "each component claims over ten thousand hours, but the entire unit only has a few thousand hours." Criterion: When reporting the MTBF of the entire unit, it must be synthesized level by level according to the BOM, and each unit's failure rate source (manual/test/field) must be labeled; if the synthesized value deviates from the measured value by more than 20%, retrace the operating condition assumptions (temperature, duty cycle, load).

4. Four Common Misunderstandings

Misunderstanding One: Treating MTBF as Lifespan. MTBF = 50,000 hours does not mean "it will not break for 5.7 years of continuous use." When committing to the lifespan of components, provide B10 or B1 lifespan and specify the confidence level, rather than giving an MTBF and letting the customer do the conversion.

Misunderstanding Two: Mixing Units. Field data is often written as "annual repair rate 1.5%," while reliability reports use FIT. Directly comparing these without conversion will inevitably lead to errors. The conversion relationship is λ(FIT) = repair rate × 10⁹ ÷ cumulative powered hours; the key is that the denominator must be powered hours rather than calendar time, otherwise, the failure rate of intermittently used products will be systematically underestimated.

Misunderstanding Three: Confusing MTTR with MTBF. Some companies include downtime due to failures in the numerator, resulting in MTBF being confused with availability A = MTBF ÷ (MTBF + MTTR). Low availability does not necessarily mean poor reliability; it could just be slow repairs — reports should provide both numbers and explain the different improvement methods.

Misunderstanding Four: Applying the Bathtub Curve Indiscriminately to Different Product Types. Modern electronic components, after thorough screening, often show a nearly flat measured failure rate curve (no early failures, no wear-out in the short term). Blindly applying the three-stage method to plan tests will waste samples; for mechanical wear components (bearings, seals, belts), the wear-out stage is obvious, and not managing life-limited components will inevitably lead to problems. The criterion comes from data, not textbooks.

5. Self-Check List

  1. □ Is the MTBF in the technical agreement clearly defined with operating conditions, statistical scope, failure definition, and confidence level requirements?
  2. □ Is the denominator of failure data the cumulative powered hours rather than the number of units or calendar time?
  3. □ When reporting MTBF, are the confidence lower limits for r = 0 and r > 0 distinguished, rather than just providing a point estimate?
  4. □ Is the current failure stage confirmed using the Weibull shape parameter or trend test, rather than relying on experience to label it?
  5. □ Is the MTBF of the entire unit synthesized level by level according to the BOM, and is the failure rate source of each unit traceable?

MTBF is the reciprocal of the failure rate, not a lifespan.

Knowledge code: 8.2.3

Version: v20261007

Author: QTank QTank is dedicated to providing systematic professional knowledge, methodologies, and practical tools for quality management practitioners, helping companies continuously improve their quality capabilities.