Measurement System Analysis (MSA) Deep Practice: From GR&R to Data-Driven Measurement System Evaluation
Measurement System Analysis (MSA) is a critical bridge in the quality management system that connects "data" with "decision-making." Whether it's a point on an SPC control chart, the calculation of process capability (Cpk), or the determination of product conformity, all quality decisions are based on measurement data. If the measurement system itself is unreliable, any quality judgment made based on data could be off by a hair and a thousand miles away from the truth.
However, in the actual operations of many manufacturing enterprises, MSA is often simplified to "conducting a GR&R once a year and submitting a report." When auditors come, they flip through the report, and if the data is acceptable, everything is fine. But the real issue is: Is the probability threshold for passing GR&R set reasonably? Does the sampling cover the entire tolerance range of the product? Are bias and linearity being overlooked? How is the long-term stability of the measurement system ensured?
This article will start from the core concepts of MSA, systematically explaining the complete methodology for measurement system evaluation, the key operational points of GR&R, the logic for determining bias and linearity, and how to build a continuous measurement system management mechanism in a digital context.
1. Core Concepts of MSA: Why Measurement Systems Need Analysis
Many quality engineers are puzzled by MSA: If the measuring instruments are calibrated, why do we still need to analyze the measurement system? The answer to this question reveals the essential difference between quality management and metrology management.
The core of metrology management is "accuracy" — whether the measuring instruments are calibrated, within their validity period, and traceable to national standards. This is a baseline, but calibration alone is far from enough. MSA focuses not only on the precision of the measuring instruments themselves but also on the variation characteristics of the entire measurement system under the influence of the "people, machine, material, method, environment" (5M) elements.
A measurement system consists of the following elements:
Measurement Equipment — gauges, inspection tools, sensors, coordinate measuring machines (CMMs), vision measurement systems, etc.
Measurement Personnel — whether the operators are trained and proficient in standard work methods.
Measurement Object — the characteristic values and tolerance ranges of the product being measured.
Measurement Method — inspection specifications, work instructions, measurement procedures.
Environmental Conditions — external factors such as temperature, humidity, vibration, and lighting.
Calibration can only ensure the accuracy of the measurement equipment under specific conditions, but it cannot cover the impact of operator differences, reading errors, fixture positioning deviations, and environmental fluctuations on the measurement results. A calibrated measuring instrument may yield vastly different measurement results when used by different operators — MSA is precisely the systematic method for quantifying these sources of variation.
The core output of MSA analysis is to determine whether the measurement system is "adequate." "Adequate" here has two layers of meaning: first, whether the total variation of the measurement system is sufficiently small relative to process variation and tolerance range; second, whether the measurement system can reliably distinguish differences between different parts.
To understand these two aspects, we need to introduce two key ratios: %GR&R and ndc (Number of Distinct Categories).
%GR&R is the percentage of the total variation (or tolerance) that is due to the repeatability and reproducibility of the measurement system. Repeatability refers to the variation when the same operator measures the same part multiple times using the same gauge, reflecting the intrinsic precision of the gauge. Reproducibility refers to the variation when different operators measure the same part using the same gauge, reflecting the impact of operator differences. GR&R is the combination of both.
ndc (Number of Distinct Categories) answers another question: How many distinct categories can this measurement system divide the process output into? If ndc is less than 2, it means the measurement system cannot reliably distinguish between "conforming" and "nonconforming" parts.
These two indicators together form the quantitative basis for MSA judgment.
2. Practical GR&R: From Scheme Design to Data Interpretation
GR&R is the most core and commonly used analysis method in MSA. However, many companies have systematic biases when executing GR&R — unreasonable sampling schemes, non-standard execution methods, and a one-size-fits-all judgment standard. This section will outline the key operational points of GR&R from a practical perspective.
1. Sampling Strategy: The Starting Point Determines the Endpoint
The sampling for GR&R is not random. Many engineers randomly pick 10 parts from the production line for analysis, resulting in a high GR&R, which is not due to a problem with the measurement system but rather because the part differences are too small. GR&R aims to assess whether the measurement system can reliably reflect differences between parts, so the sample must cover the entire tolerance range of the product — from the lower limit to the upper limit, as evenly distributed as possible.
The best practice is: Select at least 8 to 10 parts from the measurement data collected over the recent period, covering the complete range of variation from the smallest to the largest. If the process capability is very good (high Cpk), the parts produced on the production line naturally cluster around the center of the tolerance range, and it may be necessary to select or specially make parts with extreme dimensions to complete the coverage.
The automotive industry's IATF 16949 recommends selecting 5 to 10 parts, with each part measured 2 to 3 times by 3 operators. However, in practice, the sample size should be dynamically adjusted based on the risk level of the measurement system, the frequency of gauge use, and the criticality of the product. For critical safety characteristics, a combination of 10 parts × 3 operators × 3 measurements is recommended; for general process control parameters, a simplified scheme of 5 parts × 2 operators × 2 measurements is also acceptable.
2. Execution Standards: Three Common Mistakes
First, the measurement sequence must be randomized. Operators should not measure all parts in one go before the second round, nor should they measure parts in order of size. Randomization is aimed at preventing operators from forming "reading inertia" or "expected bias."
Second, the interval between measurements should simulate the actual production operation interval. Some companies have operators measure the same part 10 times in a row during GR&R, which measures "repeatability under optimal conditions" rather than actual production repeatability. The correct approach is to have operators insert other normal operations between each measurement to simulate the rhythm of daily inspections.
Third, the recording of measurement results must be confidential. Operators should not know each other's measurement results, nor should they see their own previous readings. This is a basic requirement to prevent "human correction."
3. Judgment Standards: Lower GR&R Is Not Always Better
The AIAG MSA manual (Fourth Edition) provides reference standards for GR&R judgment:
- %GR&R ≤ 10%: The measurement system is acceptable.
- 10% < %GR&R ≤ 30%: Conditionally acceptable, requiring a comprehensive judgment based on the importance of the application and the cost of improvement.
- %GR&R > 30%: The measurement system is unacceptable and must be improved.
However, there is a common misconception: The lower the GR&R, the better. This is not always the case.
When GR&R is far below 5%, it often means the measurement system is over-specified — using measuring equipment with precision far exceeding actual needs. For example, using a CMM to measure a rough machining dimension with a tolerance of ±5mm. This over-specification not only increases measurement costs but may also introduce unnecessary variation due to the high sensitivity of high-precision equipment to environmental conditions. A more reasonable approach is to choose a measurement system that is "adequate" rather than "excessive" based on the tolerance and process variation of the characteristic being measured.
Additionally, there are two main methods for calculating %GR&R: %GR&R based on total process variation (%GR&R/TV) and %GR&R based on tolerance (%GR&R/Tol). These reflect different information. %GR&R based on tolerance focuses on the ratio of measurement system variation to the tolerance band, making it more suitable for scenarios involving product conformity determination. %GR&R based on total process variation focuses on the ratio of measurement system variation to the total process fluctuation, making it more suitable for process control (such as SPC). In actual work, it is recommended to view both indicators to obtain a more comprehensive judgment.
4. What to Do If It Fails
When GR&R exceeds 30%, don't rush to replace the gauge. A more efficient method is to identify and address the sources of variation in order of magnitude.
Step one: Decompose the total variation into repeatability and reproducibility using ANOVA or the mean-range method to see which one is larger. If repeatability is dominant, the issue may be due to insufficient gauge precision, unstable measurement methods, or certain characteristics of the part being measured causing measurement difficulties (such as surface roughness leading to reading variations). If reproducibility is dominant, it indicates that the measurement results from different operators vary too much, and the focus should be on checking the consistency of operator training, the standardization of measurement methods, and whether there is room for adjustment in the fixture positioning method.
Step two: Check if the variation between parts is sufficiently large. If the part variation in the GR&R is a small proportion of the total variation, it suggests that the sample distribution range is too narrow. The solution is to supplement the sample with extreme parts covering the tolerance limits and re-analyze.
Step three: Check if the measurement system is significantly affected by environmental factors. Some precision measurements are extremely sensitive to temperature — for example, a 1°C difference between the product temperature and the ambient temperature in aluminum part measurements can cause tens of micrometers of dimensional changes. In such cases, the root cause of a high GR&R may not be the gauge but inadequate environmental control.
3. Bias, Linearity, and Stability: Three Key Dimensions Beyond GR&R
GR&R addresses the precision (Precision) of the measurement system, but accuracy (Accuracy) also needs attention. Accuracy is measured by three indicators: bias (Bias), linearity (Linearity), and stability (Stability).
1. Bias: Does the Measurement System Have Systematic Errors?
Bias refers to the difference between the observed average measurement result and the reference value (true value). The reference value is typically obtained through higher-precision measuring equipment (such as a CMM) or standard parts.
The typical process for bias analysis is: Select a standard part or calibration part with a known reference value, and have the same operator measure it at least 10 times using the gauge to be evaluated. Calculate the difference between the average measurement and the reference value. If the difference is significantly different from zero (determined by a t-test), it indicates that the measurement system has a fixed systematic bias.
The existence of bias does not necessarily mean the measurement system is unusable — the key is whether the magnitude of the bias is acceptable relative to the tolerance or process variation, and whether the bias remains consistent within the working range of the measurement system (this is what linearity analysis aims to answer).
2. Linearity: Does Bias Vary with Measured Size?
Linearity analysis addresses a deeper question: Is the bias of the gauge consistent across small and large sizes? If a micrometer has a bias of +2μm when measuring a 10mm size and a bias of +15μm when measuring a 100mm size, the bias of this measurement system is not "constant" but increases with the measured value — this is a linearity issue.
The standard method for linearity analysis is to select more than 5 standard parts covering the full range of the gauge, measure each part at least 10 times, and then perform regression analysis with the reference value as the x-axis and bias as the y-axis. If the slope of the regression line is significantly different from zero, it indicates that the measurement system has a linearity issue and needs to be corrected or the gauge replaced.
Common causes of linearity issues include: uneven wear of the gauge, design principles of the gauge leading to inconsistent precision across different ranges, or improper use (such as wear on the jaws of a vernier caliper causing increased error in large size measurements).
3. Stability: Can the Measurement System Be Consistently Reliable?
Stability analysis addresses the question of consistency over time: Are the results consistent today, next week, and next month?
The method for stability analysis is: Measure the same standard part with the same gauge under the same conditions, but distribute the measurement intervals over a longer period (several weeks to several months). By plotting the measurement results on a control chart (usually an Xbar-R chart or I-MR chart), observe if there are any abnormal trends or out-of-control points.
Stability issues are often the most overlooked in measurement system management. The calibration report shows "合格" (conformity), but the gauge may gradually deviate from the standard state due to accidental collisions, wear and tear, sensor drift, and other factors during daily use. This is why MSA is not a "one-time pass and done" task — it requires a continuous monitoring mechanism.
4. MSA Strategies for Destructive Measurement Systems
The above discussions are all about non-destructive measurements. However, in actual production, many measurements are destructive — tensile tests, hardness tests, torque tests, life tests, etc. These measurements cannot be repeated on the same part, so the standard GR&R method cannot be directly applied.
There are several alternative strategies for MSA in destructive measurement systems.
The first is nested analysis. Select multiple sets of samples from adjacent positions in the same batch (ensuring as much material consistency as possible), and use nested ANOVA to separate batch variation from measurement variation. The premise of this method is that the material uniformity of adjacent samples is sufficiently high.
The second is attribute consistency analysis. When the measurement results are binary or ordinal data (such as "合格/不合格" (conforming/nonconforming)), use a cross-tabulation Kappa coefficient to evaluate the consistency and effectiveness of the measurement system. The Kappa coefficient measures the degree of consistency between operators or between operators and the standard, after accounting for random agreement. Generally, a Kappa ≥ 0.75 indicates good consistency.
The third is production part method. For certain special destructive tests, design a set of validation parts with known results in the production process, and regularly test the reliability of the measurement system through blind testing.
The choice of method depends on the nature of the characteristic being measured, the feasibility of sample acquisition, and the risk level of the measurement system. MSA for destructive measurement systems is particularly critical in automotive component certification (such as material physical and chemical property testing) and electronic product reliability testing, but it is often omitted or simplified due to operational difficulty.
5. New Trends in MSA in the Digital Age
With the advancement of digital quality management systems, the practice of MSA is undergoing profound changes.
Automatic Data Collection and Real-Time Analysis: Traditional MSA relies on manual recording and offline calculations, which can lag and be prone to errors. In digital measurement scenarios, measurement data is automatically transmitted to the quality management system via PLC or RS-232 interfaces. The system can calculate GR&R and bias indicators in real-time and automatically trigger alerts when the measurement system shows abnormal trends. This "online MSA" model elevates the monitoring of measurement systems from an annual compliance check to a daily management activity.
Multivariate Measurement System Analysis: When measurement equipment (such as CMMs or vision measurement systems) outputs multiple characteristic measurements simultaneously, the measurement errors of these characteristics may be interrelated. Multivariate MSA uses principal component analysis or multivariate ANOVA to comprehensively evaluate the overall performance of the measurement system, identifying systemic issues that single-variable analysis cannot capture.
AI-Assisted Anomaly Detection: In the context of large volumes of automated measurement data, machine learning-based anomaly detection algorithms can help identify abnormal patterns in measurement data — for example, the bias variation trend of a gauge in a specific range, or the correlation between environmental temperature and humidity and measurement results. These tools are not meant to replace statistical methods but to assist quality engineers in quickly pinpointing areas that require in-depth analysis in a sea of data.
Digital Verification of Attribute Measurement Systems: In visual inspection and AI quality inspection scenarios, measurement results are no longer single values but composite judgments that include defect type, location, severity, and other attributes. The verification of such measurement systems cannot rely solely on traditional GR&R methods; it must combine classification model evaluation metrics such as confusion matrices and ROC curves to comprehensively assess their detection capabilities.
6. Establishing a Sustainable Measurement System Management Mechanism
Finally, upgrading MSA from a "one-time analysis action" to a "continuous management system" is a significant indicator of a company's maturity in quality management.
Graded Management: Not all measurement systems require the same level of MSA control. It is recommended to implement graded management based on characteristic classification — the measurement systems for critical safety/regulatory characteristics (CC/SC) should undergo the most frequent complete MSA; general process control characteristics should follow a simplified MSA; and non-critical reference characteristics can be indirectly verified through calibration records. The core logic of graded management is to allocate limited resources to the highest-risk measurement processes.
Periodic Review: The MSA results of a measurement system are not static. Factors such as gauge wear, personnel changes, process changes, and seasonal environmental variations can all affect the performance of the measurement system. It is recommended to establish an MSA database, analyze the historical GR&R results of each measurement system for trends, set warning thresholds, and proactively initiate improvement actions before performance indicators deteriorate but before they exceed standards.
Training and Culture Building: The reliability of a measurement system ultimately depends on the quality of the operators' execution. Regularly training inspection personnel in MSA basics, helping them understand "why MSA is necessary" and "how their operations affect measurement results," is far more effective than mandatory signatures. When frontline operators can proactively identify measurement anomalies and report improvements, the measurement system truly possesses self-healing capabilities.
Measurement System Analysis is not a one-time compliance action but the first line of defense for data-driven quality decision-making. Companies that truly understand MSA can ensure that every quality judgment is solid from the data source.
Knowledge Number: 6.2.1
Version: v20260722
Complementary Training Materials: Calibration and Gauge Management · MSA Series Integrated Practical Training (Complete PPT) — Integrates calibration traceability, MSA five characteristics and GR&R, key points of uncertainty, and digital pathways, suitable for 3-4 hours of internal training.
Complementary Tool Templates: Calibration and Gauge Management · MSA Complementary Tool Kit (Excel Templates) — Includes ledger, calibration plan, nonconformity impact assessment, periodic verification, GR&R data collection, and color-coded inspection points. Download Excel Tool Kit
Author: Quality Excellence Think Tank The Quality Excellence Think Tank is dedicated to providing systematic professional knowledge, methodologies, and practical tools for quality management practitioners, helping companies continuously improve their quality capabilities.