How Many Trials Are Enough? — A Five-Step Method for Planning DOE Trial Numbers and Power
One: 16 Trials Done, but the Conclusion is "Wasted Effort"
Engineer Xiao Lin at an electronics factory was assigned a welding process optimization task: four factors—welding temperature, welding speed, flux flow rate, and preheating time—needed to find the optimal parameter combination. At the project initiation meeting, the leader asked, "How many trials do you plan to conduct?" Xiao Lin, unsure, blurted out, "Full factorial for four factors, 16 trials." After two weeks of trials and running an ANOVA, none of the main effects of the four factors were significant. Xiao Lin scrutinized the data for three days and finally realized: it wasn't that the factors had no impact, but that the number of trials was too small—16 trials had only a 50% power to detect the effect size, and the real differences were drowned out by noise. The trials were wasted, and the production line was delayed for two weeks.
This is not an isolated case. Ask ten engineers, "How many trials are enough?" and nine won't have a clear answer, relying instead on habit and budget. The consequences of this approach are twofold: too few trials, and real effects go undetected, wasting effort; too many trials, and the budget is exceeded, leading to project cuts. Companies with tight budgets especially need to pay attention to this: when the number of trials is justified during project review, the project is more likely to be approved; if the number is vague, you will be the first to be challenged. There is a comprehensive method in DOE to answer "How many trials are enough?"—the core is trial number planning and power analysis. This article will explain it in five steps.
Two: What Determines the Number of Trials
First, let's discuss the principle. A single trial is like using a net to catch fish: if the mesh is too large, small fish (small effects) will slip through. The number of trials determines the density of this net, and it is influenced by four factors:
First, the size of the effect to be detected, δ. This is the smallest difference you want to capture, such as "for every 10°C increase in welding temperature, the welding strength must increase by 5MPa." The smaller δ is set, the more trials are needed.
Second, process noise, σ. Repeated trials under the same conditions will have inherent variations. The greater the variation, the harder it is to distinguish real differences, and the more trials are required.
Third, the two types of error risks. α is the risk of "saying there is a difference when there isn't," also known as a false positive; β is the risk of "saying there is no difference when there is," also known as a false negative. The complement of β, 1-β, is the power, or the ability to detect an effect when it truly exists—it measures the probability that the trial will detect a real difference.
Fourth, the structure of the experimental design. Full factorial, fractional factorial, adding or not adding center points, and whether to conduct replicates directly determine the total number of trials.
In summary, the number of trials is determined by δ, σ, α, β, and the design structure. δ and σ are determined by engineering reality, α and β by the risks you are willing to accept, and the design structure by the number of factors and the budget. Once you have a clear understanding of these four factors, the number of trials can be calculated without guesswork.
Three: Five-Step Method: Calculating the Number of Trials
Step One: Determine the Minimum Effect Size to Detect, δ. This step relies solely on engineering judgment, not statistical assistance. Ask yourself: how much of a parameter change is worth the investment of resources? For example, if a 5MPa increase in welding strength is necessary to justify a process change, then δ is 5MPa. Setting δ smaller than what is actually needed will result in an unnecessarily high number of trials, wasting money.
Step Two: Estimate the Noise, σ. Use historical data first: review the standard deviation from past similar trials or daily production. If historical data is unavailable, conduct 3 to 5 preliminary trials and use the variation from these to estimate σ. If time is short, estimate a slightly larger value—typically one-sixth of the tolerance band—since underestimating σ can lead to an insufficient number of trials.
Step Three: Set the Risk Levels, α and β. The industry standard is α=0.05 and β=0.20, which corresponds to 80% power. For critical characteristics involving safety or regulations, tighten β to 0.10, increasing the power to 90%, but at the cost of about a third more trials. Once these values are set, do not change them during the analysis phase.
Step Four: Calculate the Number of Trials Based on the Design Structure. This is the most mechanical step. A full factorial design with k factors is 2 to the power of k: 4 factors require 16 trials, 5 factors require 32 trials. If the number of factors exceeds 4 and the budget is limited, use a fractional factorial design: 5 factors can be screened with 8 trials. To determine if the response is curved, i.e., if the optimal value lies within the experimental range, add 3 to 5 replicates at the center point. To directly estimate experimental error, select one or two conditions for replication. Sum the design points, center points, and replicates to get the total number of trials. If a response surface optimization is planned for the next stage, axial trials belong to the second phase and should be calculated separately, not mixed together.
Step Five: Make a Decision Based on the Budget. If the calculated number of trials exceeds the budget, there are four legitimate options: increase δ to detect only larger effects; switch to a fractional factorial or screening design; accept a power below 80% but clearly state this in the report; or conduct more preliminary trials to accurately estimate σ. If none of these options are acceptable, don't do DOE; instead, perform single-factor verification. The one thing you should never do is increase α—this is trading false positives for trial numbers, with potentially serious consequences.
Returning to Xiao Lin's case: four factors, σ approximately 3MPa, and a target δ of 5MPa. With 16 full factorial trials, the power was only about 50%, clearly insufficient. By increasing δ to 8MPa, meaning an 8MPa increase in strength is necessary to justify a process change, the same 16 trials could achieve a power of over 80%. The issue wasn't the number of trials but the overly aggressive setting of δ. Adjusting δ would have made the conclusions robust.
Four: The Four Most Common Pitfalls
Pitfall One: Following Habit for Trial Numbers. "We did 8 trials last time, so we'll do 8 this time"—the number of factors, σ, and δ have all changed, so how can the number of trials remain the same? Recalculate the number of trials before each round.
Pitfall Two: Treating α Like Play-Doh. If the results are not significant, adjust α from 0.05 to 0.10 to force a "significant" result. This is self-deception and can be easily spotted by audit experts.
Pitfall Three: Reporting Only p-Values, Not Power. p-Values only answer whether there is a difference; power answers how large a difference can be detected. Stating in the report that "the power to detect effects of 5MPa or more is 85%" is more convincing than listing ten p-values.
Pitfall Four: Calculating Trial Numbers After the Trials. Realizing after the trials that the effects cannot be detected means starting over. Trial number planning must be completed during the experimental design phase, documented in the trial plan, and archived along with the basis for δ and σ values.
Trial number planning is not complex, but it determines whether DOE is "getting it right the first time" or "wasting a round." Next time you are asked "how many trials," don't guess; lay out the values of δ, σ, α, and β, and the answer will become clear.
The number of trials is not a guess but is calculated based on δ, σ, α, β, and the design structure.
Knowledge Number: 6.4.1
Version: v20260831
Author: Quality Excellence Think Tank Quality Excellence Think Tank is dedicated to providing systematic professional knowledge, methodologies, and practical tools for quality management practitioners, helping companies continuously improve their quality capabilities.