Advancing QE Skills (29) | Practical Weibull Analysis: Predicting Failures Using Lifespan Data

By: QTank Published: 10/9/2026 Views: 13
Current rating: ★★★☆☆ Rate this Equivalent to 8 ratings

1. Handing Over 5 Failure Points to Software Yields a "Lifespan Conclusion"

A batch of electric actuators from a manufacturing company began to experience early sticking in the field. After-sales requested the quality department to provide a "reliable lifespan of this batch of products." The QE's approach was to record the failure times of the 5 repaired products into the software, select "Weibull two-parameter," and obtain a shape parameter β (beta) = 1.6, characteristic life η (eta) = 4200 hours, and B10 life = 1300 hours. The report stated, "The product's B10 life is 1300 hours, and it is recommended to replace every 1000 hours." During the review, several issues were exposed: first, the 5 items belonged to two completely different failure modes—3 were due to seal aging, and 2 were due to gear wear, which were forced into a single curve; second, the 30 items that were still functioning normally at 4000 hours (censored data) were not recorded, and the software only used the failure points; third, the report did not mention what mechanism β = 1.6 corresponds to, how reliable this estimate is, or the width of the B10 confidence interval. The result was that the curve could neither explain the past nor predict the future. Weibull analysis is not difficult to calculate, but it is challenging to classify the data, handle censored data, interpret the parameters, and determine how far the conclusions can be extrapolated.

2. Key Principles: What Each of the Three Parameters Controls

The core of the Weibull distribution is its three parameters. Understanding them is essential to make sense of the curve, which is more than just a string of numbers.

Shape Parameter β—Controls "Failure Mechanism." β is the most informative number, directly indicating the current stage of the product's lifespan:

Range of β Failure Type Engineering Implications and Countermeasures
β < 1 Early Failure (Decreasing Failure Rate) Process defects, inconsistent incoming materials, assembly damage; rely on screening and aging to eliminate, not on extending lifespan
β ≈ 1 Random Failure (Constant Failure Rate) External forces/occasional impacts; at this point, the exponential distribution and MTBF are meaningful
β ≈ 1.5~2.5 Early Wear, Partial Fatigue Bearing, seal, and contact wear; clear wear trends
β ≈ 3~4 Typical Wear Failure Material fatigue, accelerated wear; this is the most effective interval for preventive replacement
β > 4 Rapid Wear or Existence of a Lifespan Threshold Significant location parameter γ (gamma); a three-parameter Weibull model should be used

Scale Parameter η—Controls "Level." η is the characteristic life, the point at which exactly 63.2% of the products have failed. It serves as the scale for shifting the curve left or right without changing its shape.

Location Parameter γ—Controls "Minimum Life." γ is significant only when no samples fail before a certain time point. Blindly fitting a three-parameter model can make the fit "look better" but lose its mechanistic meaning. Generally, start with γ = 0 and use goodness-of-fit tests to determine if it needs to be introduced.

B10 Life is the most frequently cited metric in engineering—the time point at which 10% of the products fail. When γ = 0:

B10 = η × (0.10536)^(1/β)

Some common values: when β = 1, B10 ≈ 0.105η; when β = 1.5, ≈ 0.223η; when β = 2, ≈ 0.325η; when β = 3, ≈ 0.472η; when β = 3.44 (approximate normal distribution), ≈ 0.520η. Note that the smaller β is, the earlier B10 is relative to η—this corresponds to the phenomenon of "the first batch of products failing first" when early failures dominate.

3. Five Practical Steps: From Raw Records to Deliverable Conclusions

Step One: Classify the Data, Not Rush to Fit. First, label each record as "failure" or "censored (still functioning at the time point)," then group them by failure location and failure mode. Criterion: Only allow separate modeling for a group if the number of failure samples r ≥ 5; if r < 5, merge similar mechanisms or provide a qualitative description. If there are two or more failure modes and the modes vary with stress, separate curves must be built, and mixed fitting is prohibited.

Step Two: Perform Probability Plotting/Goodness-of-Fit Testing. Plot the data for the same mechanism in the median rank coordinate F(i) = (i − 0.3) ÷ (n + 0.4) based on the failure order i, and check if they form an approximate straight line. Criterion: A correlation coefficient r ≥ 0.95 from least squares fitting is acceptable; if the points show a clear upward or downward curvature, it indicates mixed mechanisms or the need for a three-parameter model, not just "the data is like this." For stricter tests, use the Anderson-Darling test, where a p-value of at least 0.05 is required.

Step Three: Select Estimation Method, Correctly Handle Censored Data. Use maximum likelihood estimation (MLE) when there is a significant amount of censored data, as it can utilize the time information of all non-failed samples; use rank regression when data is scarce and almost all samples have failed. Criterion: The number of censored samples should not exceed 60% of the total sample size; directly deleting non-failed samples is the most fatal data error in Weibull analysis. Key Reminder: If a curve has no failure points (all censored), β cannot be estimated, and any "lifespan prediction" is purely speculative.

Step Four: Interpret β, Calculate B10, and Provide Confidence Intervals. First, interpret the mechanism corresponding to β based on the table from Step Two, and confirm it aligns with the on-site disassembly conclusions; if it does not, question the data grouping before questioning the software. Then, use the formula to calculate B10. Criterion: The 90% confidence intervals for both β and B10 must be provided, and it should be specified whether they are one-sided lower limits or two-sided intervals; if the interval width exceeds 30% (e.g., B10 = 1300 hours with an interval of 900 to 2600 hours), the conclusion can only be described as a "reference value" and cannot be used as a design input for replacement cycles.

Step Five: Validate Sample Size Using Target Back-Calculation. When there is no historical data and new tests are required for validation, the most common method is the zero-failure scheme, with a one-sided reliability confidence lower limit of (1 − C)^(1/n). To ensure this lower limit is at least 0.90 (i.e., to prove B10 meets the standard), at least 22 samples must each run to the target B10 time without failure at 90% confidence; 29 samples are needed at 95% confidence. This conclusion is independent of β and is very practical—it directly tells you how far "all 3 samples passing" is from "validating B10."

4. Four Common Misconceptions

Misconception One: Mixing Different Failure Modes into a Single Curve. Seal aging is dominated by corrosion/material degradation, while gear wear is due to contact fatigue; both have different β and η values. A mixed fit β will fall between the two, overestimating early failure risks and underestimating the wear stage's acceleration, leading to a completely distorted B10. Correct Approach: First, disassemble the failed items to determine the mode, then fit the data in groups.

Misconception Two: Treating Censored Data as Noise and Discarding It. Recording only failure points is equivalent to retaining only the "worst samples," systematically underestimating η and B10. The software must record the operational times of non-failed products as suspended/censored data.

Misconception Three: Focusing Only on B10, Ignoring the Confidence Interval of β. The estimation of β itself has uncertainty, and the fewer the samples, the wider the interval. A β = 1.6 estimated from 5 failure points may have a 90% confidence interval as wide as 0.9 to 2.8—spanning from "early failure" to "wear," making the mechanistic judgment unreliable. Suggested Phrasing: Report "β = 1.6 (90% CI: 1.1 to 2.3), corresponding to a wear failure tendency," rather than just "β = 1.6."

Misconception Four: Extrapolating Lifespan Beyond the Data by One or Two Orders of Magnitude. The reliability of Weibull extrapolation decreases sharply with the extrapolation distance. Extrapolating B10 from 2000 hours to 20000 hours is equivalent to magnifying the model's error by ten times. Empirical Boundary: B10 should not exceed three times the longest test duration; beyond this range, it can only be called a "trend judgment," not a "lifespan conclusion."

5. Self-Check List

  1. □ Have the data been classified into failure and censored categories and grouped by the same failure mechanism (with at least 5 failure samples per group)?
  2. □ Has a probability plot linearity or Anderson-Darling goodness-of-fit test been performed (with r ≥ 0.95)?
  3. □ Have all censored data been fully recorded, rather than directly deleted?
  4. □ Have the 90% confidence intervals for both β and B10 been provided, and has the corresponding failure mechanism for β been interpreted?
  5. □ Is the extrapolation distance for B10 within three times the longest test duration? If new tests are required for validation, has the sample size been calculated based on the target (zero-failure scheme at 90% confidence ≥ 22 samples)?

A curve that does not distinguish mechanisms is wrong, no matter how straight it is.

Knowledge code: 8.2.3

Version: v20261009

Author: QTank QTank is dedicated to providing quality management professionals with systematic knowledge, methodologies, and practical tools to continuously enhance corporate quality capabilities.