QE Skill Enhancement (10) | What's the Difference Between Cp/Cpk and Pp/Ppk: Correct Interpretation of Short-term and Long-term Capability
A precision machining company supplies drive shafts to a new energy client. During the client's audit, a process capability report was required. Using the same production line, the same batch of 200 shaft diameter data points, and the same software, two engineers came up with different conclusions: one reported a Cpk of 1.46, concluding "adequate capability"; the other reported a Cpk of 1.21, concluding "improvement needed." After verifying the formulas, data, tolerances, and measurement tools, they found no discrepancies except for how the subgroups were defined—former grouped every 5 consecutive pieces, while the latter grouped by shift. The situation became even more challenging when the client's SQE casually calculated a Ppk of 1.10 using the entire dataset and returned both reports. The company's confusion was very typical: why would a single process capability have three different numbers, and which one is the correct one?
1. Key Principles: Two Different σ for Two Different Variations
The formula structures for Cp, Cpk, Pp, and Ppk are identical, differing only in the denominator used for standard deviation:
- Within-group (short-term) σ: This is estimated using the variation within subgroups, commonly using the range method R̄/d2 or the standard deviation method S̄/c4, representing the "instantaneous variation of the process at the moment."
- Overall (long-term) σ: This is calculated using the sample standard deviation s of all data, including both within-group and between-group variations.
Statistically, the relationship is σ²overall ≈ σ²between-groups + σ²within-groups. This means that the difference between Ppk and Cpk almost entirely comes from the "between-group" variation—shifts, material batches, tool wear, changeovers, environmental factors, multi-cavity molds, and parallel machines. Cpk answers "what the capability would be if production were only within the selected window," while Ppk answers "what the client will actually see in the batch of goods received." The former represents potential capability, and the latter represents actual performance, which is the true meaning of the short-term and long-term distinction.
This explains the case above: σ within-group is not an inherent property of the data; it depends on how you define the subgroups. Grouping every 5 consecutive pieces excludes slow drifts from the within-group variation, resulting in the smallest σ and the highest Cpk. Grouping by shift includes within-shift drifts, increasing σ and decreasing Cpk. The larger the subgroup, the closer the within-group σ is to the overall σ, and the closer Cpk is to Ppk. In the extreme case, if the entire dataset is treated as one subgroup, the "calculated Cpk" numerically equals Ppk. Ppk has only one calculation method, making it the only number that can be compared across reports—this is why clients only recognize Ppk during the initial process study phase.
2. Practical Steps: Five Steps to Convert Capability into Decision-making Conclusions
Step 1: Determine stability before calculating capability. Use data from a controlled window to draw a control chart. The criteria for stability are: no out-of-control points in 25 consecutive subgroups, no 8 points on the same side of the center line, no 6 points in a row increasing or decreasing, and no 2/3 points outside the 2σ limits. Cpk calculated before stability is confirmed is only the instantaneous variation of a specific window and should not be used externally; clients always see Ppk.
Step 2: Define subgroups based on the source of variation, not the number of data rows. Subgroups should only include common causes: continuous outputs from the same shift, the same material batch, and the same tool cycle. Subgroup size should be 4 to 10 (recommended 5), with no fewer than 25 subgroups and a total sample size of at least 100 pieces (125 pieces or more for key characteristics). Multi-cavity molds, multi-station equipment, and parallel machines must be studied separately by cavity, station, and machine. This subgroup definition rule should be included in the report—part of the conclusion, not just a format.
Step 3: Calculate both sets of indices in one go. Use the same dataset to calculate Cp/Cpk (within-group σ, S̄/c4 or R̄/d2) and Pp/Ppk (overall s). If the software defaults to "capability analysis" and the wrong method is selected, it might use the overall σ in the Cpk formula or vice versa, both of which are common errors that do not trigger an error message but produce a seemingly reasonable number.
Step 4: Use two ratios for diagnosis, not just the index values.
- Ppk/Cpk ≥ 0.85: Between-group drift can be ignored, the process is basically stable, and improvements can be made based on the Cpk value.
- Ppk/Cpk between 0.70 and 0.85: There is identifiable system drift, prioritize investigating shift differences, material batch differences, tool and mold wear, and first article inspection after changeovers.
- Ppk/Cpk < 0.70: The process is not yet stable, do not report Cpk externally, stabilize the process first, then discuss capability.
- Cpk/Cp ≥ 0.90: The center offset is within 0.3σ, the improvement direction is to reduce variation.
- Cpk/Cp < 0.80: The center offset exceeds 0.6σ, adjust the center first—move the mean back to the center of the tolerance, which is usually much cheaper than reducing variation.
Step 5: Choose the appropriate index based on the object and provide a range. For initial process studies, PPAP, and new project releases, report Ppk, with the reference threshold being ≥ 1.33 for general characteristics and ≥ 1.67 for key and safety characteristics (based on the client's CSR). For ongoing production, follow the CSR requirements; if the client requires both indices, report both and specify the σ method and subgroup structure. The capability index itself has sampling error: when the true value is 1.33, the 95% confidence interval half-width for n=100 is approximately ±0.20, and for n=30, it expands to ±0.36 or more. When the sample size is less than 50 pieces, the report should provide the lower confidence limit or directly use the lower confidence limit for judgment, rather than signing off on a point estimate. Recalculation triggers should be included in the control plan: re-stabilize after the control chart goes out of control, change material batches or tools, change equipment and fixtures, change process parameters or drawings, and regular re-evaluation every 3 months.
3. Common Misconceptions
Misconception 1: Using Excel's STDEV to calculate Cpk and R̄/d2 to report Ppk. The former treats the overall σ as the within-group σ, systematically underestimating Cpk and misguiding the improvement direction to "too much variation"; the latter treats the within-group σ as the overall σ, overestimating Ppk, leading to reports that cannot stand up to client complaints.
Misconception 2: Mechanically dividing subgroups by "every 5 rows as one group." Grouping without considering the source of variation is like treating randomly combined data as a homogeneous window: it can either include data from different shifts, masking drifts, or split data that should be in the same group, inflating the within-group σ. The numbers come out quickly, but the conclusions are entirely a matter of luck.
Misconception 3: Treating "Cpk and Ppk values being the same" as evidence of perfect process control. If only one subgroup is defined, or if the subgroup is so large that it approaches the entire sample, the two values will inevitably be the same. This is a result of the algorithm's degradation, not an excellent process. When the two values are identical, the first reaction should be to verify the subgroup structure, not to write "process is stable and controlled."
Misconception 4: Claiming Cpk ≥ 1.67 with a sample size of n=30 or even n=10. The index itself fluctuates greatly with small samples, and a point estimate meeting the criteria does not mean the capability meets the criteria. Without specifying the sample size, confidence interval, and judgment criteria, the report lacks the authority to release the process.
Misconception 5: Always attributing Ppk being lower than Cpk to calculation errors or fraud. It is a direct evidence of between-group drift, and the correct action is to investigate shifts, material batches, tools, and changeovers, not to modify the criteria to cover up the issue. Conversely, a high Cpk does not necessarily mean good client experience—the client's experience is described by Ppk.
4. Self-check List
- Capability calculation was performed after using a control chart to determine stability, and the stability conclusion (25 subgroups, no out-of-control points, no chains, no trends) is archived with the report.
- The report specifies the subgroup size, number of subgroups, total sample size, and σ calculation method (R̄/d2 or S̄/c4 and overall s), allowing others to replicate the calculation.
- All four indices—Cp, Cpk, Pp, and Ppk—are provided, and the two ratios Ppk/Cpk and Cpk/Cp have been interpreted for diagnostic purposes.
- Ppk is reported externally, and Cpk is used internally, consistent with the client's CSR; if the sample size is less than 50 pieces, provide the lower confidence limit.
- Recalculation triggers (after re-stabilization, changing material batches or tools, changing equipment and fixtures, changing drawings or parameters, and regular re-evaluation) are included in the control plan.
Capability indices are not just arithmetic problems but statements about "where the variation comes from." Clarifying the σ method, explaining the subgroup structure, and attributing the differences to specific sources can turn the awkward situation of three different numbers from the same dataset into an actionable improvement directive.
The difference is not in the formula, but in where σ comes from.
Knowledge code: 6.3.2
Version: v20260920
Author: QTank QTank is dedicated to providing systematic professional knowledge, methodologies, and practical tools for quality management practitioners, helping companies continuously enhance their quality capabilities.