Six Sigma Project Stalled in the M Stage? —— A Case Study of Data Collection Plan Rebuilding in an Automotive Parts Company
1. D Stage Passed Smoothly, but M Stage Nearly Derailed the Project
An automotive parts company launched a Six Sigma project last year: the first-time pass rate on the welding line had been hovering around 92% for a long time, and the customer had complained about insufficient welding strength for three consecutive quarters, explicitly requiring a 60% reduction in PPM by the end of the year. The project team, led by a Black Belt, performed excellently in the D stage — the project charter, SIPOC, VOC conversion, and the definition and target of Y (reducing the welding defect rate from 8.2% to below 2%) passed the review in one go, boosting the management's confidence.
However, once the project entered the M stage, it hit a snag. Following the "convention," the Black Belt created an Excel defect record form and sent it to three shifts, asking the inspectors to fill in each welding defect daily. Two weeks later, when the data was collected, the project team held a regular meeting and became increasingly uneasy:
- The three shifts had completely different standards for "false welding": the morning shift relied on visual appearance, the afternoon shift used sound from tapping, and the night shift only discovered it during the tensile test.
- The morning shift filled in over 800 records, while the night shift had only 60 — the night shift inspectors were busy with shipping and filled the form in a rush during the last ten minutes before their shift ended.
- The same type of defect appeared with four different scrap codes in the system, with "porosity" alone being recorded as "porosity," "weld hole," "sand eye," and "pinhole."
- More troubling was that when the same inspector retested the same batch of products, the inconsistency rate between the two judgments exceeded 20%.
At the third project meeting, management asked only one question: "Do you believe these data yourselves?" The Black Belt was silent. The project was almost halted.
This scenario is all too common in Six Sigma projects. Among the five stages of DMAIC, the M (Measure, measurement) stage is the least valued but is the most vulnerable to being "fatally flawed by unusable data." Many projects do not fail in the analysis stage but rather in the M stage — if the data is garbage, even the most beautiful A, I, and C stages are like building on sand.
2. What Exactly Should the M Stage Deliver?
First, let's align on the concept. The M stage is not just about "collecting data," but delivering four key items:
Operational Definition: What exactly is Y, how is it judged, and what are the criteria? Without an operational definition, "defects" are a murky account.
Measurement System Verification: Is the "ruler" used to measure the data accurate? Visual judgment, gauges, and testing equipment must all undergo measurement system analysis (MSA). If the repeatability and reproducibility (GR&R) are unsatisfactory, the data is not credible.
Data Collection Plan: What data to collect, from where, how much, how, and who is responsible. This includes sampling methods, stratification dimensions, sample size, record forms, and review mechanisms.
Baseline and Process Capability: On the basis of reliable data, calculate the current level — defect rate, sigma level (Z value), and process capability index — to serve as the starting point for subsequent improvements.
The case company had none of these four items, which naturally made the M stage difficult to advance.
3. Four Pitfalls in the Case, Each Making the Data Unusable
Pitfall One: No operational definition, "defects" are understood differently. Is "false welding" an appearance issue or a strength issue? How much can be seen visually? Does tapping sound count? The project team had not even agreed on what constitutes a defect, so the data collected by the three shifts was not on the same scale, making the aggregated defect rate meaningless.
Pitfall Two: Measurement system not verified, the ruler itself is skewed. The inconsistency rate in retesting exceeded 20%, indicating extremely poor repeatability in the inspectors' visual judgments. Data measured with such a "ruler" is mixed with a large amount of measurement error, leading to distorted conclusions in hypothesis testing and regression analysis.
Pitfall Three: Sampling based on "convenience" rather than design. Relying solely on the inspectors' convenience, the morning shift recorded more data, while the night shift recorded less. The most severe problems were often missed during the night shift. The sample lacked representativeness, making the calculated defect rate a "selected truth."
Pitfall Four: Recording mechanism failed, no one was responsible for the data. The form design was unreasonable, filling it out was voluntary, there was no review, and the data was scattered in each person's own Excel file. When it came time to analyze, missing fields and incorrect codes were discovered, and it was too late to correct them.
4. Five Steps to Rebuild the Data Collection Plan
With the assistance of external consultants, the project team started over and used five steps to solidify the M stage.
Step One: Write operational definitions to a testable level. The team broke down "welding defects" into four categories: false welding, porosity, spatter, and weld-through, providing quantitative criteria for each. For example, "false welding = weld point separation area exceeds 30% or tensile strength is less than 120N." Standard photos and physical samples were attached and hung at the inspection stations. New inspectors judged according to the samples, and the criteria for the three shifts were finally unified.
Step Two: Fix the ruler before measuring the data. A GR&R analysis was conducted on the visual judgment, and the results were indeed unacceptable. The team provided three days of training for the inspectors, unified the criteria, and retested, reducing the GR&R to 18%, an acceptable level. The tensile testing equipment was recalibrated, and the calibration cycle was shortened from one year to six months.
Step Three: Design stratified sampling, not "whoever has time fills it out." Stratified sampling was designed based on shifts, equipment, and material batches. Each shift was required to randomly select 30 finished products daily for comprehensive testing, using a random number table to decide which items to test. The sample size was calculated to ensure that a 1% difference could be detected, ensuring that the data was both representative of the whole and sufficient for analysis.
Step Four: Standardize the recording and review mechanisms. A one-page record form was redesigned, with defect types as drop-down options to eliminate the "four different ways of writing." Each inspected sample was photographed for evidence. The team leader reviewed the day's records before the end of each shift, and the quality engineer conducted weekly spot checks on data completeness. Data now had someone responsible.
Step Five: Pilot the new mechanism for three days before formal implementation. The new mechanism was piloted for three days to check the data completeness rate, coding standardization rate, and judgment consistency. Only after all criteria were met did the team formally collect data for two weeks to establish the baseline. The results showed a welding defect rate of 8.2%, with a process Z value of about 1.4, which was consistent with the estimates from the D stage. Management finally acknowledged the data as usable.
5. With a Solid Data Foundation, the Rest of the Project Proceeded Smoothly
After establishing the baseline, the project entered the A stage almost effortlessly: a stratified Pareto chart analysis revealed that porosity defects accounted for 62% of all defects and were highly concentrated in the night shift of Welding Machine No. 2. Combining this with regression analysis, the team identified two key factors: low protective gas flow and excessive wire extension. In the I stage, the process parameters were adjusted, and a gas flow monitoring system was installed. In the C stage, control charts were used to verify the effectiveness and update the standard work instructions. Six months later, the welding defect rate dropped from 8.2% to 1.1%, and the PPM at the customer site decreased by over 70%, with a financial benefit of approximately 1.8 million yuan, and the project was closed on schedule.
Looking back, the true determinant of the project's success or failure was not the sophisticated statistical tools used in the A stage, but the solid data foundation laid in the M stage. Four key takeaways are worth remembering:
First, Measure is not just "collecting data," but "designing data." Data does not fall from the sky; it is "designed" through operational definitions, measurement systems, and sampling plans. The more detailed the design, the easier the subsequent analysis.
Second, data quality takes precedence over data quantity. A thousand records with inconsistent criteria are not as useful as three hundred records with consistent criteria. The primary goal of the M stage is to "make the data credible," not to "make the data abundant."
Third, the measurement system is the ruler of the data; an inaccurate ruler is useless. No matter how beautiful the statistical methods, they cannot save data with excessive measurement errors. If the GR&R is unsatisfactory, fix the ruler before discussing analysis.
Fourth, a slow M stage leads to a fast A, I, C stage. Many teams rush through the M stage to get to the analysis, only to repeatedly return to collect more data, making the process even slower. Slowness is speed, the most straightforward principle in Six Sigma projects.
6. In a Nutshell
The M stage is the foundation of DMAIC: the four deliverables — operational definition, measurement system, sampling design, and baseline — are all essential. With a solid data foundation, the project will proceed smoothly.
Data is not collected, it is designed — establish a solid baseline in the M stage, and the rest of the Six Sigma project can proceed.
Knowledge code: 6.1.1
Version: v20260818
Author: Quality Think Tank Quality Think Tank is dedicated to providing systematic professional knowledge, methodologies, and practical tools to quality management practitioners, helping companies continuously improve their quality capabilities.