QE Skill Advancement (16) | Introduction to DOE: Transitioning from Single-Factor Experiments to Full-Factorial Experiments
1. Adjusting Parameters for Three Months, but the Conclusion Fails on the Production Line
A manufacturing company was working on improving the flatness of stamped parts. The engineer, following the old practice of "changing one parameter at a time," first fixed the pressure at 12 MPa and adjusted the speed from 30 to 45, reducing the flatness from 0.28 mm to 0.24 mm, and concluded that "speed is effective." Next, with the speed fixed at 45, the pressure was adjusted from 12 to 14 MPa, further reducing the flatness to 0.22 mm, and concluded that "pressure is also effective." After three months, the report concluded that "both speed and pressure are significant factors, and it is recommended to set the speed at 45 and the pressure at 14." However, when the production line followed these recommendations, the flatness returned to 0.27 mm, and the improvement effects were almost entirely lost.
Upon reviewing the process, the issue was identified: the speed of 45 was effective only under the condition of a relatively low pressure of 12 MPa, which suppressed the springback. Once the pressure was increased to 14 MPa, the benefit of this combination disappeared. The "speed 45 + pressure 14" setting was never actually tested in the single-factor experiment design, so this scenario was never observed. Similar situations frequently occur in the daily work of quality engineers (QEs): the "significant factors" in the reports often fail in the field, not because the data is inaccurate, but because the experimental method can only answer single-variable questions, while the field is always a multi-variable problem.
2. The Essence of the Transition: From "Conditional Effects" to "Effects + Error"
In a single-factor (OFAT) experiment, the readings actually contain three components: the effect of the target factor, the interaction between factors, and pure noise. The design structure of OFAT cannot separate these components, leading to four statistical consequences.
First, there is no error degree of freedom. Each round changes only one factor and measures only one result, attributing all variations to that factor without an independent error estimate. The denominator for t-tests or F-tests does not exist, and significance can only be judged by "the value seems to have decreased significantly." The criterion is simple: an error degree of freedom (df_error) of ≥ 3 is the minimum threshold for勉强能判, and ≥ 6 is considered stable, while the typical df_error for an OFAT experiment is 0.
Second, the estimated effects are conditional, not main effects. When factors A and B interact, the effect of A measured at a fixed level of B is the main effect of A plus a portion of the interaction effect that is folded in. Changing the fixed level of B will change the measured value. This manifests in the field as "this parameter sometimes works, sometimes it doesn't," and is eventually recorded as "experience."
Third, experimental points are scattered along the coordinate axes, wasting information. Each "new experiment" in OFAT must be repeated under a different condition of another factor, starting from the same point. In contrast, full-factorial experiments place the experimental points at the vertices and center of a hypercube, fully covering all combinations of factor levels. Each column has a balanced number of positive and negative levels (orthogonal), and the estimates of main effects and interactions of various orders are not mixed.
Fourth, two key diagnostic capabilities are added. One is the curvature test: if the mean response at the center point significantly differs from the mean response at the corner points (p < 0.05), it indicates that the response surface is curved, and two levels are insufficient; the experiment should then transition to a response surface design. The second is the lack-of-fit test: it determines whether the model truly represents the data or is just a forced fit.
For comparison, using the same budget: a 2-factor, 2-level experiment with 3 center points each, conducted about 20 to 21 times, can estimate the main effects, second-order interactions, curvature terms, and 3 pure error degrees of freedom. The same number of OFAT experiments can only provide two uncorroborated conditional effects, plus a vague statement that "the overall variation is large."
3. Five Practical Steps to Transition from OFAT to Full-Factorial Experiments
Step 1: Confirm that the response can be used for judgment. The response must be a measurable and repeatable quantity, and the measurement system must be validated first. Criteria: for key characteristics, GR&R (P/TV) ≤ 10% can be directly used for DOE; 10% to 30% can be used, but the report must note the impact of measurement uncertainty on the effect size; > 30% requires repairing the measuring instrument, otherwise the experiment will be dominated by noise.
Step 2: Reduce the number of factors to 2-4 and set levels within the range of field variations. Use a cause-and-effect matrix or fishbone diagram to score and reduce the candidate factors. Do not force a full-factorial design if the number of factors exceeds 5 (screen first). Set levels within the range that can be achieved in production, not at the equipment limits. The distance between two levels should be ≥ 3 times the equipment resolution; otherwise, the set values themselves cannot be distinguished.
Step 3: Determine the scale, ensuring sufficient error degrees of freedom before starting. Design a 2^k full-factorial experiment (k = 2 to 4) with 3 to 5 center points. If high noise is expected, repeat the corner points twice (doubling the number of experiments). Three hard criteria: error degrees of freedom ≥ 3 (recommended ≥ 6); for the smallest engineering effect size you want to capture, statistical power ≥ 0.8 (calculate using software, not post hoc); at least 3 center points are needed to perform the curvature test.
Step 4: Randomize the experimental order. Randomize the run order to avoid synchronization of equipment warming, tool wear, shift differences, and factor levels. When noise sources are obvious, group by shift or batch, randomize within groups, and analyze between groups. Record the actual levels of each run, the operator, batch number, and environmental conditions; do not substitute "set values" for "actual values."
Step 5: Analyze based on criteria, not feelings. To determine significance, both conditions must be met: p < 0.05, and the effect size reaches the engineering threshold (e.g., flatness 0.03 mm, which is 10% of the tolerance). If only the former is met, record it as "statistically significant but engineering negligible." When interactions are significant, do not report only the main effects; perform simple effects analysis (fix one factor at each level and compare the effects of the other factor) and illustrate the interaction direction. Model validation involves three checks: adjusted R² ≥ 0.80, lack-of-fit p > 0.05, and no trends or funnel shapes in the residual plots. Models that do not meet these criteria can only be used to indicate direction and should not be used to set parameters. Finally, conduct 3 to 5 confirmation tests with the optimal combination, and the actual mean must fall within the 95% confidence interval of the predicted value; otherwise, return to step 3 to recheck noise and the model.
If the analysis yields more than one optimal combination (e.g., two feasible combinations with similar effects), prioritize the one that is less sensitive to noise and has a wider operating window, which is more cost-effective than conducting another round of experiments.
4. Five Common Misconceptions
Misconception 1: Using "change one, observe once" data to judge significance. Without error degrees of freedom, noise is mistaken for effects, and the direction of amplification or reduction is entirely random.
Misconception 2: Treating conditional effects as main effects. Changing the fixed level of a factor can reverse the conclusion, leading to the belief that "DOE is unreliable" — the unreliable part is the design, not the method.
Misconception 3: Starting full-factorial experiments without replication or center points. In a 2^4 design without replication, the four main effects and six second-order interactions consume all degrees of freedom, leaving zero error degrees of freedom. Higher-order interactions are forced to act as errors, leading to inflated p-values and unverifiable conclusions.
Misconception 4: Setting parameters independently based on main effects even when interactions are significant. The report may provide "speed 45, pressure 14" as two independent optimal values, but the product's flatness can vary significantly across other combinations of these two intervals because the true effect is the combination itself.
Misconception 5: Treating p-values as the only standard. Significance only indicates that "this difference is unlikely to be random noise," not that the difference is large enough to warrant changing the production line. Whether to change parameters also depends on passing the engineering threshold and cost considerations.
5. Self-Check List
- The measurement system for the response has been validated: key characteristics GR&R ≤ 10%, general characteristics ≤ 30%, data archived
- Factors are 2 to 4, levels are within the range of field variations, and the distance between two levels is ≥ 3 times the equipment resolution
- Power has been calculated before starting: for the target effect size, power ≥ 0.8; error degrees of freedom ≥ 3 (recommended ≥ 6), source clearly identified (center points or repeated corner points)
- The experimental order has been randomized and, if necessary, grouped; actual levels and noise conditions for each run have been recorded
- When analyzing, if interactions are significant, perform simple effects analysis first; the model meets R²(adj) ≥ 0.80, lack-of-fit p > 0.05, and 3 to 5 confirmation tests have been conducted
- Conclusions have been converted into "parameter combination rules" in the control plan and work instructions, rather than two unrelated intervals
Single-factor experiments show effects, but full-factorial experiments reveal the truth.
Knowledge code: 6.4.1
Version: v20260926
Author: QTank QTank is dedicated to providing systematic professional knowledge, methodologies, and practical tools for quality management practitioners, helping companies continuously improve their quality capabilities.