Advancing QE Skills (9) | Measurement System Analysis for Destructive Testing: Nested Design and Alternatives

By: QTank Published: 9/19/2026 Views: 16
Current rating: ★★★☆☆ Rate this Equivalent to 8 ratings

A modified plastic company supplies interior components to an original equipment manufacturer (OEM). Tensile strength is a critical indicator for both incoming and finished product quality control. During the PPAP phase of a new project, the customer requested a measurement system analysis report for tensile strength. The laboratory's approach was to take 10 consecutive injection-molded test strips from the same batch, have the same tester break each one in sequence, and use the range of these 10 results to estimate repeatability. They then divided this by the total variation, resulting in a %GRR of only 8.6%, and concluded that the "measurement system is excellent." The customer's SQE dismissed this with one sentence: These 10 test strips are 10 different test objects. The 8.6% you calculated includes both the non-uniformity of the test strips themselves and the measurement error, and you have not separated these two factors—so does it represent the measurement system or the material's non-uniformity? Even more embarrassing for the company: when another tester recalculated using their own data, they got 19%; when a different testing machine was used, it became 27%. For the same measurement system, the conclusions ranged from "excellent" to "marginal," and no one dared to sign off on it. The root of the problem is not incorrect calculations but the wrong design structure—destructive testing cannot simply follow the cross-sectional GR&R design.

1. Key Principles: What Nested Design Can and Cannot Separate

The foundation of conventional GR&R is "the same object can be measured repeatedly": several operators measure the same batch of parts in rotation, with the parts remaining unchanged, allowing the variation to be cleanly separated into "between parts" and "measurement system" components. Destructive testing removes this foundation—once a test strip is broken or a sample is crushed, it is gone, and the same test unit can only be measured once.

Therefore, a nested design must be used: "a group of units that are almost homogeneous" replaces "the same unit," with units nested under operators and groups. There is a crucial statistical fact to recognize: within the same group, each unit is a different physical object. The variation between units mathematically equals "true between-unit variation + measurement variation," and these two components are not separable (lack of identifiability) in the data structure of a single measurement. The nested design is valid in engineering based on a replacement assumption—that the units within a group are sufficiently homogeneous, so that the residual variation within the group approximates pure measurement variation.

From this, three inferences can be drawn:

  • The stricter the grouping, the more reliable the conclusion, but the narrower the process range it covers. Compressing the production time window to 30 minutes and the same mold cavity results in the highest homogeneity within the group, but at this point, the "between-part variation" component is almost eliminated, making the %GRR overly optimistic. Relaxing the grouping to span shifts increases the coverage range, but the homogeneity within the group collapses, and the residual variation is contaminated by between-unit differences, leading to an inflated %GRR and potentially misjudging good equipment as faulty. This is a trade-off that must be explicitly stated and cannot be had both ways.
  • Low degrees of freedom, wide confidence intervals. Under the nested structure, each operator measures each group only once, and the number of units used to estimate the residual variation is limited. The relative uncertainty of the %GRR point estimate is typically on the order of 30%. Therefore, the upper confidence limit, not the point estimate, should be used for decision-making.
  • Variation structure is more valuable than the total number. The proportions of between-group variation, operator variation, equipment variation, and residual variation directly determine whether to repair the fixture, standardize the SOP, or first control the incoming material non-uniformity.

2. Practical Steps: Five Steps to Complete MSA for Destructive Testing

Step 1: Define the structure and fix the "group" definition. The grouping rules must be traceable: the same equipment, the same shift, and adjacent units produced consecutively, with every 3 units grouped together; units within the same group must not span different mold cavities, batches, or time windows (the time window should not exceed 30 minutes). Design scale: at least 2 pieces of equipment × 2-3 testers × 12-20 groups, with 3 units per group. Each tester measures one unit from each group (units within the group are not reused). The total number of measurements should be no less than 40-60 times, and for key characteristics, it is recommended to have more than 60 measurements.

Step 2: Use indirect characteristics to pre-verify "within-group homogeneity." This is the only self-check foundation of the nested design. The criteria must be written into the plan: the range within the same group of units for non-destructive characteristics strongly related to the destructive indicator (such as size, density, hardness, weight, etc.) should not exceed 1.5 times the 6σ measurement variation of that characteristic's measurement system; or the standard deviation within the group should not exceed one-third of the standard deviation between groups. If this is not met, reduce the time window, regroup by mold cavity or batch, and do not proceed with contaminated groups—this is the most common entry point for nested design failure.

Step 3: Use the sandwich method to define repeatability (core alternative criterion for non-retestable scenarios). Since separation is not possible, use the sandwich method:

  • Lower Bound: Perform 25 repeated measurements using a standard test block or standard force gauge to get σ_low, representing the basic repeatability of the equipment and operator; or independently estimate using a repeatable non-destructive measurement step on the same sample (such as measuring size before breaking).
  • Upper Bound: Consider the variation between units within a homogeneous group as "measurement variation + residual between-unit variation." Since between-unit variation is non-negative, its standard deviation σ_high is necessarily the upper bound of repeatability.
  • Judgment Rule: 6σ_low / T ≤ 10% and 6σ_high / T ≤ 30% can directly determine that the measurement system is acceptable (a conservative conclusion holds); 6σ_low / T ≤ 10% but 6σ_high / T exceeds 30% should not be directly judged as "nonconforming." First, check the within-group homogeneity and process dispersion—these are two completely different causes; if 6σ_low / T already exceeds 10%, the measurement system itself is nonconforming, and there is no need to proceed with the nested design. The report must state that "repeatability falls within the X% to Y% range," not just a single number.

Step 4: Nested ANOVA to produce a variation component table. Decompose the between-group, operator, equipment, and residual components according to the nested model, and calculate %GRR as the sum of operator variation and residual variation. The denominator must be clearly specified: use the tolerance band width T as the denominator, or use the total variation TV as the denominator according to customer requirements, and report both. Judgment: ≤10% is acceptable; 10% to 30% is conditionally acceptable, and a protection band (commonly 3σ measurement variation within the judgment limit) must be set and the risk explained; >30% is not acceptable. Note the use of the upper confidence limit for judgment: a point estimate of 15% may already have an upper 95% confidence limit of 30% in a nested design; do not claim "repeatability is less than 10%" if there are fewer than 12 groups.

Step 5: Complete bias, linearity, and long-term stability. Bias is verified using standard test blocks or by comparing the same batch of test pieces with a third-party laboratory, with the criteria |bias| ≤ 10% of the tolerance band width and a single-sample t-test showing no significant difference; laboratory comparison can use |En| ≤ 1. Linearity should be tested at least at low, medium, and high levels, with the bias change across the range not exceeding 5% to 10% of the tolerance band width. Stability is maintained by retaining a "mother piece" or standard block, retesting it monthly or every 500 tests, and plotting a control chart; after equipment maintenance, fixture replacement, tester change, or changes in procedures or methods, the critical steps must be redone.

3. Common Pitfalls

Pitfall 1: Using the range of adjacent samples as repeatability. This confuses the differences between test objects and measurement errors. The direction of the conclusion's bias depends on the within-group homogeneity, which can be positive or negative and is unpredictable. Moreover, it often only covers a single operator and a single machine, completely missing the reproducibility variation.

Pitfall 2: Reporting only a single number, ignoring variance components and degrees of freedom. The sample size and number of repetitions in destructive testing are naturally limited by cost, leading to a wide confidence interval for the %GRR point estimate. Not reporting the interval, not calculating the degrees of freedom, and not considering the proportion of components results in a report that cannot be used for decision-making.

Pitfall 3: Grouping across time periods, mold cavities, and batches. Different homogeneity within groups can inflate the residual variation, leading to an inflated %GRR. The final conclusion is "equipment is nonconforming," but the actual issue is the grouping method. Replacing the equipment will not solve the problem.

Pitfall 4: Assuming automatic data recording eliminates human influence. The largest variation in destructive testing often comes from sample preparation, holding length, sampling position, loading rate, and fixture wear, all of which are related to "how people do it." Operator variation is part of the measurement system and must be included in the model.

Pitfall 5: Quietly changing the denominator. Using 6σ / T to determine conformity and 6σ / TV to report to the customer can result in numbers that differ by several times. The denominator used and the reason for its use must be clearly stated on the first page of the report.

Pitfall 6: Completing the MSA once and archiving it. Not redoing the MSA after changing raw materials, testers, or major equipment repairs means that the %GRR several months later is no longer the same as the initial number. Process variations will record these drifts, leading to misjudgments that the process itself is deteriorating.

4. Self-Check List

  • Grouping rules are traceable: same equipment, same shift, consecutive production, clear time window, no spanning of mold cavities or batches within the group
  • Quantitative evidence of within-group homogeneity (indirect characteristic range or standard deviation compared between groups) is documented and archived
  • The report provides an interval conclusion for repeatability (σ_low to σ_high), not an isolated %GRR number
  • The variance component table is complete, including between-group, operator, equipment, and residual variations, and it specifies the direction for corrective actions
  • The upper confidence limit is used for judgment, and the denominator (T or TV) is clearly stated; the protection band for the 10% to 30% interval is included in the control plan
  • The conditions for retesting bias, linearity, and stability after changes are written into the laboratory management procedures

Destructive test samples can only be measured once, but their reliability can still be clearly calculated. Acknowledge the premise that "the same sample cannot be tested again," and use the design structure, within-group homogeneity, and interval conclusions to clarify the inseparable parts—a honest interval estimate is more effective in maintaining the baseline for release decisions than a single, pretty number.


If retesting is not possible, use intervals rather than point values.

Knowledge code: 6.2.1

Version: v20260919

Author: QTank QTank is dedicated to providing systematic professional knowledge, methodologies, and practical tools for quality management practitioners, helping companies continuously improve their quality capabilities.