GR&R Only 8%, but Customer Retest Exceeds Limits? —— Five-Step Sampling Plan for MSA

By: QTank Published: 10/8/2026 Views: 16
Current rating: ★★★☆☆ Rate this Equivalent to 8 ratings

A quality engineer at a die-casting company conducted a gauge repeatability and reproducibility (GR&R) study, calculating a %GRR of 8.2%, which is considered "acceptable" according to the common AIAG criteria. Three months later, a customer conducted a process audit using the same caliper, the same batch of products, and the same set of forms, but the result was 31%, deemed "unacceptable," and a nonconformity was issued. The engineer compared the two reports line by line: the gauge hadn't changed, the formula was the same, and the personnel were the same two individuals. The only difference was that the customer randomly selected 10 pieces from the production line, while the engineer had picked 10 pieces from the warehouse on that day.

The problem lies in this last sentence. GR&R is not a standalone tool that can independently "measure" the quality of a gauge; its conclusion is highly dependent on one thing—what parts you measure.

1. The Denominator of %GRR is What You Put In

The standard formula for %GRR is the measurement system variation (GRR) divided by the total variation (TV), which includes repeatability, reproducibility, and part-to-part variation (PV). In other words, the denominator is not fixed; it is determined by the sampling.

This has a subtle consequence: if the parts sampled have a large variation (from different batches, different cavities, or even intentionally selected to include large, medium, and small sizes), the part-to-part variation is increased, the denominator becomes larger, and the %GRR is "diluted" to a very small value, making the report look particularly good. When the customer retests using actual production parts, the part variation returns to a normal level, the numerator remains unchanged, and the %GRR immediately increases.

This is the most common root cause of "internally excellent, but exceeding customer limits": it's not that the gauge suddenly became bad, but that the denominators of the two studies are fundamentally different. Therefore, the first step in MSA is not to do the math, but to align the sampling range with the actual variation in daily production.

Using a set of illustrative numbers makes this clearer. Assume the variation introduced by the gauge (GRR) is 0.03mm, then:

  • When the part-to-part variation in the sample is 0.30mm, the total variation is approximately 0.30mm, and the %GRR is about 10%, which is judged as "acceptable";
  • When the part-to-part variation in the sample is 0.10mm, the total variation is approximately 0.10mm, and the %GRR is about 30%, which is judged as "unacceptable".

The same gauge, the same operator, the same formula, but the conclusion changes from "acceptable" to "unacceptable" simply because the dispersion of the parts is different. This also explains a common phenomenon: the truly difficult parts to measure in the workshop are often not those with large errors but those with very stable processes and minimal part variation—these have a naturally higher %GRR and require finer resolution and stricter measurement procedures, rather than a more expensive gauge.

2. Five-Step Sampling Plan

Step One: Define the sampling range, i.e., the boundary of "process variation."

The sampling range should either cover the actual variation range of the process or the product specification range (use the specification range if the process capability is insufficient), taking the wider of the two. The basis for this should not be subjective but should rely on existing data: previous control charts, process capability studies, and recent inspection records. Once the range is defined, it should be written into the MSA plan, and the report should clearly state "what range these 10 pieces represent."

Step Two: Select parts—cover the range, don't deliberately pick extremes.

Sample 10 pieces from daily production at regular time intervals (e.g., one piece every half hour), allowing the sample to naturally include the normal variation of the process, including pieces close to the upper and lower limits. Three prohibitions:

  • Do not take all pieces from the same batch, the same cavity, or the same shift, as this will make the part-to-part variation approach zero, inflating the %GRR;
  • Do not artificially select extreme pieces or extend the range beyond daily production, as this will dilute the %GRR to a falsely low value;
  • Do not use first articles, prototype pieces, or selected pieces as samples, as they do not represent mass production.

Parts should be renumbered (e.g., randomly labeled 1 to 10), and the original numbers should not be exposed at the site.

Why 10 pieces? Because in a single study, you need to estimate three components: repeatability, reproducibility, and part-to-part variation. Too few samples will make the part-to-part variation estimate unstable, leading to random conclusions; too many samples will significantly increase the number of measurements, making it difficult to coordinate on-site. Ten pieces strike a balance between precision and cost. If it is indeed impossible to gather 10 conforming pieces (e.g., small batches, expensive pieces), a nested design with fewer operators can be used, but this must be clearly noted and not used to draw a conforming conclusion based on "insufficient data."

Step Three: Select operators and trial times—use the usual people, in the usual way.

Operators should be the actual inspectors or operators who use the gauge daily. Select 2 to 3 individuals, and do not substitute engineers or the "most skilled one." Each operator should measure each piece 2 to 3 times, with a typical configuration being 10 pieces × 3 operators × 3 times = 90 measurements. Measuring each piece only once by each operator will not allow for the separation of repeatability, making the study conclusions invalid.

Step Four: Randomization and blind testing—skip this step, and the first three are nullified.

  • Randomize the measurement sequence, not following the part numbers or a fixed operator sequence;
  • Operators should not see the original part numbers or their previous readings;
  • Zeroing, clamping, and reading the gauge should all follow the usual procedures, without any additional steps to "improve accuracy";
  • Multiple measurements by the same operator should not be consecutive on the same piece, allowing time for memory effects to decay.

Step Five: Judgment and documentation—consider both %GRR and ndc.

The criteria for judgment involve two indicators: %GRR is typically required to be less than 10% (acceptable), 10% to 30% is decided based on the importance of the application and the cost of improvement, and greater than 30% is unacceptable; simultaneously, the number of distinct categories (ndc) = PV/GRR should be at least 5. Both should be reported. The report must clearly state the source of the sample, the covered range, and the configuration of parts, personnel, and measurement times, otherwise, the customer cannot determine whether this study represents their usage scenario.

3. Three Most Common Sampling Errors

Error One: Adjusting the sample to "get good results." This is the most damaging to credibility. The purpose of GR&R is to expose problems, not to pass audits; once the sample is adjusted, the numbers may look good but lose their decision-making value, and the customer's retest will inevitably reveal the discrepancy.

Error Two: Parts are too similar, leading to a false nonconformity. Contrary to dilution, if the 10 pieces are almost identical (e.g., taken consecutively from the same box), the part-to-part variation is extremely small, and the %GRR will be higher, leading to a false judgment of the gauge as nonconforming and triggering an unnecessary gauge replacement. Both biases point to the same issue—sampling range error, making the conclusion invalid.

Error Three: Not repeating the study after changing materials, molds, or shifts. The process variation range changes, so the denominator changes, and the old GR&R is no longer valid. After any significant changes in process, materials, measurement procedures, or gauges, the sampling plan should be re-planned and the study repeated.

4. One-Sentence Summary

The denominator of %GRR is determined by the sampling; if the sampling range is wrong, even the most beautiful conclusions cannot stand—first, ensure the sample represents the actual production variation, then discuss whether the gauge is conforming or nonconforming.


The conclusion of the measurement system depends on what parts you measure.

Knowledge code: 6.2.1

Version: v20261008

Author: QTank QTank is dedicated to providing systematic professional knowledge, methodologies, and practical tools for quality management practitioners, helping companies continuously improve their quality capabilities.