Day Shift Release, Night Shift Return? —— Counting MSA Transforms Visual Inspection from "By Feel" to "By Standard"
1. Introduction: The Same Product, Qualified in the Morning, Unqualified in the Afternoon
A certain automotive parts company manufactures the outer shell of a new energy battery pack connector, and visual inspection is the final checkpoint before shipment. However, in the first quarter of this year, the customer returned six batches of goods, all for the same reason: "Visual defects missed during inspection." The workshop was not convinced: the shells undergo three rounds of full inspection after production, so how could any defects be missed? The quality department brought the returned samples back to the factory for re-inspection, and the results were even more awkward—on the same shell, the day shift inspector deemed it "qualified" and released it, while the night shift inspector glanced at it and said, "Scratch, return." Both felt they were correct.
Scenes like this play out almost every day in this company: Inspector Zhang, with eight years of experience, can identify scratches with a single glance; new inspector Li, who is unsure about the vague "slight color difference," can only ask the team leader. Customer complaints, internal rework, and departmental disputes all stem from one root cause: The "standard" for visual inspection is not in the measuring tools but in each person's mind.
Visual inspection is a typical counting measurement system—judgment results are binary conclusions of "qualified/unqualified" rather than continuous numerical values. There are no objective readings like a caliper; it all depends on the human eye and judgment. Once the judgment criteria become ambiguous, the measurement system fails. This article will detail the complete improvement process of this company, explaining how to conduct a counting MSA, what problems it can solve, and the most common pitfalls.
2. Why Visual Inspection is Most Easily Influenced by "People"
First, let's distinguish between two types of measurement systems: measurement systems use numerical values, such as calipers measuring diameters and three-coordinate measuring machines measuring position accuracy. The main sources of variation are the measuring tools themselves and the operation methods, analyzed using GR&R. Counting measurement systems, on the other hand, use "yes or no" to make judgments, such as visual inspection, go/no-go gauges, and air-tightness tests. The primary source of variation is "people"—defect definitions, experience, environment, and condition.
The sources of variation in counting measurement systems typically include five aspects:
- Ambiguous Definitions: The procedure document states "no obvious scratches allowed," but how "obvious" is that? Length, width, location, and quantity are not quantified.
- Lack of Limit Samples: Without physical standards for the "boundary between qualified and unqualified," edge cases rely solely on personal judgment.
- Environmental Differences: Lighting brightness, observation angles, and workstation distances are not standardized, causing the same defect to appear differently under different lighting conditions.
- Experience Differences: Experienced employees rely on their experience, while new employees rely on their courage, leading to vastly different judgments on the same product.
- Fatigue and Condition: Visual inspection heavily depends on attention. After two hours of continuous inspection, the misjudgment rate visibly increases.
These variations compound, resulting in: The measurement system itself "produces" nonconforming products—either by missing them, leading to customer complaints, or by over-rejecting, leading to internal rework. The most terrifying aspect of this problem is that, unlike dimensional deviations, there is no data alarm; it hides in every "I think it's okay."
3. Case Analysis: A Counting MSA Transforms "Feel" into "Data"
The company decided to stop the finger-pointing and speak with data. The quality department led a complete counting MSA, following a straightforward process.
Step One: Prepare Samples. From the inventory of the past three months, 50 shells were selected: 20 clearly qualified, 20 clearly unqualified, and 10 "edge cases"—samples that are on the boundary between qualified and unqualified, testing the inspectors' judgment. All samples were renumbered to hide their original judgments, and the "standard judgments" were determined by process and quality engineers as the baseline.
Step Two: Blind Testing. Three inspectors (an eight-year veteran, a two-year experienced hand, and a three-month newcomer) independently judged the samples in two rounds, with a day's interval between the rounds, and the sample order was randomized. The judgment results were compared with the baseline, and the consistency between inspectors was also calculated.
Step Three: Calculate Kappa. Kappa is a common indicator for measuring consistency, ranging from -1 to 1. The main difference from simple consistency rates is that it deducts the component of "random guessing"—two people can guess correctly half the time, but Kappa measures the consistency after excluding luck. Industry reference: a Kappa of 0.75 or above is considered good, 0.40 to 0.75 is moderate, and below 0.40 indicates poor consistency.
The results were clear: The overall Kappa was only 0.31, indicating poor consistency; out of the 10 edge cases, only 3 were judged consistently by all three inspectors. The details were even more disheartening: among the 50 samples, the veteran's consistency with the baseline was 92%, while the newcomer's was only 74%; and between the veteran and the newcomer, 11 products were judged differently—6 of which were edge cases.
The data clearly highlighted three issues: First, there were no limit samples, so edge cases relied entirely on personal judgment; second, defect definitions were vague terms like "slight" and "obvious," making them unenforceable; third, there was a gap in experience, with the judgment rules existing only in the veteran's mind and not documented as organizational knowledge.
4. Targeted Solutions: Four Steps to Improve Consistency from 0.31 to 0.87
With the issues clear, the improvements were targeted. The company took four steps.
Step One: Create Limit Samples and Standard Photos. For the four types of defects—scratches, color differences, burrs, and dirt—limit samples were created for the "upper limit of qualification"—anything lighter is qualified, anything heavier is unqualified. High-definition standard photos were also provided. The limit samples were confirmed by the customer representative, process, and quality engineers, locked in the sample cabinet at the inspection station, and registered with regular calibration. From then on, "slight scratch" was no longer a vague term but a specific criterion: "visible scratch length greater than 3mm and located on surface A is deemed unqualified."
Step Two: Write a Defect Dictionary and Judgment Rule Cards. The quantified rules for the four types of defects were made into A4 cards and posted at each inspection station: defect type, quantified indicators, judgment conclusion, and handling method. The cards also specified lighting requirements (illumination no less than 800 lux) and observation posture, standardizing the environmental differences.
Step Three: Training and Certification. The three inspectors conducted blind testing exercises using the limit samples, with immediate feedback and explanations for incorrect judgments, repeated over five rounds. After the exercises, they were re-evaluated, and only those who achieved a Kappa score could work independently. The newcomer, Li, was re-evaluated two weeks later, during which time the veteran provided guidance and reviewed each piece.
Step Four: Establish a Regular Review Mechanism. The company stipulated: a small-scale blind test (20 samples) to review Kappa every quarter; new employees must pass the consistency test before starting work; limit samples must be re-confirmed by the three parties every six months to prevent "standard drift."
5. Results and Reflections
Three months after the improvements, the results of the second complete counting MSA were as follows: the overall Kappa improved from 0.31 to 0.87, with all three inspectors achieving over 90% consistency with the baseline, and the consistency rate for edge cases increased from 30% to 85%. More importantly, the business results were significant: customer complaints about visual defects dropped from 6 per quarter to 1; internal inspection disputes decreased by about 80%, and rework and scrap due to misjudgment were significantly reduced; the time for new employees to work independently was shortened from three months to one and a half months.
Three lessons were learned from the reflection:
First, the samples must include "edge cases." If only clearly qualified and clearly unqualified samples are used, the Kappa score will be artificially high, and the real issues will not be exposed. Edge cases are the "mirror" of the measurement system.
Second, achieving a Kappa score is not the end. Personnel turnover, product changes, and wear on limit samples can all cause consistency to degrade quietly. The company later found that after a surface treatment process adjustment, the consistency in color difference judgments dropped significantly, and the quarterly review mechanism promptly identified the issue—reviews are not just formalities but are like "health checks" for the measurement system. In terms of quality costs, the investment of a half-day's work for a quarterly review is negligible compared to the cost of a customer complaint due to a missed defect, which could result in the entire batch being returned, plus shipping costs and production line stoppages. This is a cost-effective investment.
Third, standards must be "visual and comparable." No matter how rigorous the textual descriptions are, they are not as effective as a limit sample. Moving the standards from the file cabinet to the workstation ensures that consistency moves from "the veteran's mind" to "the company's asset."
6. One Sentence Summary
The quality of visual inspection fundamentally depends on the consistency of judgment; the value of a counting MSA is not just in calculating a Kappa score but in exposing vague standards and solidifying them with limit samples and quantified rules—when inspection no longer relies on a specific individual, quality is truly established.
The greatest variation in counting measurement systems comes from "people": using limit samples, defect dictionaries, training and certification, and regular reviews to solidify judgment standards, visual inspection can transform from "by feel" to "by standard."
Knowledge code: 6.2.1
Version: v20260824
Author: Quality Think Tank Quality Think Tank is dedicated to providing systematic professional knowledge, methodologies, and practical tools for quality management practitioners, helping companies continuously improve their quality capabilities.