QE Skill Advancement (5) | MSA for Attribute Gauges: How to Quantify Visual Inspection (Kappa and Consistency Analysis)
1. Introduction: Why a "Beautiful" Kappa Report Was Rejected by the Customer
A certain electronics manufacturing company supplies products to an overseas customer, with appearance requirements of "no visible scratches, no color differences." During a supplier process audit, the customer issued a nonconformity: the lack of an attribute consistency analysis report. The quality engineer quickly completed and submitted the report: 3 inspectors, 50 samples, each making two rounds of judgments, with an overall Kappa of 0.81 and a consistency rate of 94%, concluding that "the measurement system is acceptable."
Two weeks later, the report was returned. The customer's SQE asked only one question: "How many of these 50 samples were 'difficult to judge' marginal items?" Upon review, it was found that 46 were clearly conforming items, 4 were clearly nonconforming items, and there were zero marginal items. In other words, the inspectors did not need to use their judgment to provide answers—high consistency and high Kappa were simply the result of "too easy questions."
This scenario highlights the core of attribute MSA: it is not just a statistical calculation, but a test of the measurement system's capability. The reliability of the conclusion depends first on the design of the sample and criteria, and then on the Kappa value. The basic content (how to calculate Kappa, how to build limit samples) has been covered. This article delves deeper: how to proportion samples, what each of the three consistency criteria measures, how to set acceptance thresholds, how to draw conclusions, and how to close the loop on nonconformities.
2. Key Principles: What Exactly Does Attribute MSA Measure?
Components of Variation. The judgment results of an attribute measurement system can be broken down into four parts: the true classification of the product (baseline), the repeatability variation of the inspector, the reproducibility variation between inspectors, and the ambiguity of the baseline itself. The first two parts are separated through experimental design, while the third part is quantified and eliminated through the definition and use of limit samples.
Division of Consistency Rate and Kappa. Consistency rate = number of consistent judgments ÷ total number of judgments, which is easy to calculate but includes the component of "guessing correctly." Even with random guessing, the long-term consistency rate for a binary classification can be around 50%; the more skewed the ratio between the two categories, the higher the chance consistency rate—when all items are conforming, anyone judging all as "conforming" can achieve a consistency rate of over 90%. Kappa = (Po - Pe) ÷ (1 - Pe), where Po is the observed consistency rate and Pe is the chance consistency rate, it subtracts the contribution of "guessing correctly." This is why a single sample composition can result in a high consistency rate and a high Kappa, but the measurement system actually proves nothing.
Three Consistency Criteria, Each Measuring Something Different.
- Inspector Self-Consistency (Within Appraiser): Whether the same person's judgments are stable over two or more rounds, reflecting repeatability and answering "Will today's judgment be the same tomorrow?"
- Consistency of Each Inspector with the Baseline (Each Appraiser vs Standard): Reflecting accuracy (bias) and answering "Is the judgment correct?"
- Consistency Between Inspectors (Between Appraisers): Reflecting reproducibility and answering "Do different people use the same standard?" Note its trap: if two people judge according to the same incorrect standard, the consistency rate can be perfect, so it can only be used as an auxiliary criterion and not as the sole basis for acceptance.
Miss Rate and False Alarm Rate, More Business-Relevant Than Kappa.
- Miss Rate = number of times a nonconforming item is judged as conforming ÷ total number of nonconforming items, directly determining the risk of nonconforming items flowing out.
- False Alarm Rate = number of times a conforming item is judged as nonconforming ÷ total number of conforming items, directly impacting internal rework, scrap, and production capacity.
The cost implications of these two rates are opposite, and their tolerances differ. It is entirely possible to have a situation where "Kappa meets the standard, but the miss rate is 8%," meaning that while the overall consistency is good, the risk is primarily on the customer's side—so all three indicators must be reported together, not just one.
3. Practical Steps: From Scheme Design to Conclusion Implementation
Step 1: Make the Baseline Objective. The standard judgment for attribute MSA must be objectively verifiable: for items that can be quantified using measurement tools (scratch length, color difference ΔE, gap size), measure the values first and then set limits; for items that cannot be quantified, the process, quality, and customer (or design) teams must sign off on the limit samples and clearly state "if it is worse than this, it is judged as nonconforming." Any samples with an indeterminate baseline must be excluded and not included in the analysis. At the same time, standardize the lighting (generally no less than 800 lux, color temperature 5000~6500K), observation distance and angle, and observation time per item (e.g., no more than 15 seconds).
Step 2: Proportion Samples to Expose Issues.
- Total sample size ≥ 50 items; nonconforming items should account for 30%~50% to ensure sufficient observations for both miss and false alarm rates.
- Marginal items (near the boundary between conforming and nonconforming, typically within ±10% of the limit value) should be at least 20%, i.e., at least 10~12 items.
- Conceal sample numbers, mix batches, and interleave conforming and nonconforming items. For quarterly reviews, the sample size can be reduced to 20~30 items, but the proportion of marginal items must not be reduced, otherwise the conclusion is invalid.
Step 3: Design the Experiment. Inspectors ≥ 3, covering three levels: experienced, average, and new employees; each inspector makes 2~3 rounds of judgments, with intervals between rounds ≥ 4 hours or overnight, and the sample order is randomized within each round; blind testing, no communication between inspectors, and no visibility of the baseline judgment. This method is applicable when the product can be re-introduced and re-judged; for destructive appearance judgments (such as opening or disassembling), use limit sample comparison or nested schemes.
Step 4: Draw Conclusions Based on the Criteria Table. Common thresholds (follow the customer's stricter requirements if applicable):
- Self-consistency ≥ 90% and Kappa ≥ 0.75 is good; 80%~90%, Kappa 0.6~0.75 is marginal, requiring rectification; below 80% or 0.6 is unacceptable.
- Consistency with the baseline: all inspectors ≥ 90%, and each individual no less than 85%; if any individual is below 85%, they should undergo retraining.
- Consistency between inspectors: Kappa ≥ 0.75, used only as an auxiliary criterion.
- Miss rate ≤ 5%, safety/critical characteristics should be controlled to ≤ 2%.
- False alarm rate ≤ 10%, which can be adjusted based on production line rhythm and rework costs.
Additionally, the 95% confidence interval for Kappa and the sample composition must be provided. When n = 30, the interval width often exceeds 0.2; if the lower limit is below 0.6, even if the point estimate meets the standard, it should be judged as "insufficient evidence, need to expand the sample"—this is why many reports fail to stand up to customer scrutiny.
Step 5: Handling and Closing the Loop. The conclusion should not just state "acceptable/unacceptable," but must include three elements: retraining and re-evaluation arrangements for nonconforming personnel (using new, untrained samples, with the same standards as before); conditions for re-running the MSA (before new employees work independently, product changeover, update of defect standards or limit samples, changes in lighting or workstations, and quarterly regular reviews); and record retention (sample list and baseline judgments, original judgment forms, statistical reports, and tripartite sign-off records) for customer audit and traceability.
4. Common Pitfalls
- Using Consistency Rate as the Sole Criterion. When 90% of the items are conforming, the consistency rate is naturally high, but it does not reveal true judgment ability. Consistency rate, Kappa, miss rate, and false alarm rate should all be reported.
- Not Including Marginal Items in the Sample, or Using Items Judged as Conforming by Everyone. Kappa can be artificially inflated, and the sample composition will be exposed by the customer; marginal items are the most informative part of the entire analysis.
- Only Calculating Consistency Between Inspectors. Even if multiple people make the same incorrect judgment, the consistency rate can still be perfect, and the product will still be misjudged and flow out; consistency with the baseline is the criterion for accuracy.
- Only Conducting One Round of Judgments. Without repeated rounds, it is impossible to separate self-repeatability variation; only the correctness of a single judgment can be seen, not the stability of the inspector.
- Treat Kappa 0.75 as a Hard Threshold. Kappa is highly influenced by the ratio of the two sample categories; the same 85% consistency rate can result in a Kappa difference of over 0.3 under different baseline ratios; only providing the point estimate without the confidence interval and sample composition makes the conclusion unverifiable.
5. Self-Check List
- Each sample has an objective basis for standard judgment (numerical judgment or tripartite sign-off on limit samples), and ambiguous items have been excluded.
- Sample size ≥ 50 items, nonconforming items account for 30%~50%, and marginal items ≥ 20%.
- ≥ 3 inspectors, 2~3 rounds, intervals between rounds, random order, blind testing without communication.
- The report includes self-consistency, consistency with the baseline, consistency between inspectors, miss rate, false alarm rate, and provides the Kappa confidence interval and sample composition.
- Nonconforming personnel have retraining/re-evaluation records, and the conditions for re-running the MSA and record retention are clearly defined.
6. One-Sentence Summary
The reliability of attribute MSA comes from the design of the sample and criteria, not just the Kappa value.
Without marginal items in the sample, a high Kappa is just a result of the questions being too easy.
Knowledge code: 6.2.1
Version: v20260915
Author: QTank QTank is dedicated to providing quality management professionals with systematic knowledge, methodologies, and practical tools to help companies continuously improve their quality capabilities.