Experiment Completed, but the Wrong Conclusion Was Chosen? — A Comprehensive Case Study of DOE Result Analysis in an Electronics Company

By: QTank Published: 8/25/2026 Views: 36
Current rating: ★★★☆☆ Rate this Equivalent to 8 ratings

Many companies, after introducing Design of Experiments (DOE), focus all their efforts on "designing the plan": how many factors to choose, how to set the levels, whether to use full factorial or fractional factorial designs, and so on, meticulously deliberating. However, when the experiments are completed and the data is in front of them, they often only calculate the average values, compare highs and lows, and hastily conclude. The result is often: the experiment was conducted, money was spent, but the wrong conclusion was chosen. This article uses a case study of wave soldering process improvement in an electronics manufacturing company to thoroughly review a scenario where "the experiment was done, but the data was not properly analyzed," and how the correct DOE result analysis should be conducted.

1. Case Background: 18 Experiments Completed, but Yield Remained Unchanged

A certain electronics manufacturing company primarily produces power boards, with a long-standing wave soldering defect rate of around 2.8%, where three types of defects—cold soldering, bridging, and solder shorts—account for 80% of the issues. Customers have complained for two consecutive quarters, and the company has pledged to reduce the defect rate to below 1.0% within three months.

The process team selected four candidate factors for a full factorial experiment: preheating temperature (100~140°C), solder pot temperature (245~265°C), conveyor chain speed (0.8~1.4 meters/minute), and flux spray volume (0.6~1.0 milliliters/minute). They used a 2^4 full factorial design with an additional 2 center points, totaling 18 experiments, each randomly sampling 200 boards to calculate the defect rate.

The experiments were completed in three days, but the data left the team perplexed. Some team members compared the average defect rates of the four factors individually: the average defect rate was 1.6% at high solder pot temperature and 3.0% at low temperature, leading to the conclusion that "the higher the solder temperature, the better"; the average defect rate was 2.3% at both high and low chain speeds, leading to the conclusion that "chain speed has no effect"; preheating temperature and spray volume also had "no effect." Therefore, the team set the solder pot temperature to the upper limit of 265°C and the chain speed to the high speed of 1.4 meters/minute—justifying it with "since there's no effect, running faster can increase production capacity."

A week of trial production, however, delivered a harsh reality: the defect rate increased to around 2.5%, far from the 1.0% target; and the trial records clearly showed a defect rate of 0.7%, which could not be replicated. Some team members began to murmur, "Is DOE just a matter of luck?"

2. The First Hurdle: Averaging Values, Averaging Out Interaction Effects

The problem lies in the analysis method. "Comparing the average values of individual factors" has a fatal premise—that the factors do not interact with each other. Once there are interaction effects between the factors, this method will average out the truth.

Consider the combination data of solder pot temperature (B) and chain speed (C) in this experiment: the defect rate was 2.1% at 245°C and slow chain speed, 3.9% at 245°C and fast chain speed; 0.7% at 265°C and slow chain speed, and 2.5% at 265°C and fast chain speed. The effect of chain speed is entirely dependent on the solder temperature: at high solder temperature, a slow chain speed is necessary, while a fast chain speed is worse. Averaging the four combinations results in a defect rate of 2.3% for both fast and slow chain speeds, leading to the conclusion that "chain speed has no effect"—not that it has no effect, but that the effect is averaged out.

Interaction effects are very common in engineering: higher temperatures improve solder flow, and the board needs to stay in the oven longer (slow chain speed) to fully wet. When the temperature is insufficient, staying longer increases oxidation. The two factors must be considered together, not individually.

3. The Correct Approach: Look at the Model First, Then the Effects

The correct DOE analysis is not about averaging values but conducting Analysis of Variance (ANOVA). The process involves four steps.

Step 1: Examine the overall model. First, check if the model's P-value is significant and if the lack-of-fit term is not significant, confirming that the experiment was not in vain and that the model is not missing any terms. Step 2: Test each term. List the main effects and all second-order interaction terms, and check their P-values one by one, eliminating non-significant terms. In this case, the solder pot temperature (P < 0.001) and the "solder pot temperature × chain speed" interaction term (P < 0.001) are significant, while the other terms are not. Step 3: Check the goodness of fit. The adjusted R-squared value reaches 0.93, indicating that the model explains most of the variability. Step 4, and the most crucial rule: if an interaction term is significant, the main effect cannot be interpreted independently. The statement "the higher the solder temperature, the better" is only valid when the chain speed is fixed.

Many teams stumble at this point: they focus solely on the P-values of the main effects, either ignoring or misunderstanding the interaction terms, and end up using a "half-baked model" to adjust parameters, inevitably leading to failure.

4. Interaction Plots: Translating Statistical Conclusions into Engineering Language

To make statistical results actionable, visual aids are essential. Plot the interaction between factors B and C: the x-axis represents chain speed, and the two lines represent high and low solder temperatures. If the lines are almost parallel, it indicates no interaction; if they cross, it indicates strong interaction. In this case, the two lines form a clear "X" shape—this is the most直观 (intuitive) evidence of interaction.

From the plot, the engineering conclusion can be read: at a solder temperature of 265°C, a slow chain speed of 0.8 meters/minute is required, with an expected defect rate of about 0.7%; at a solder temperature of 265°C with a fast chain speed, the expected defect rate is 2.5%. Combining the main effect plots, interaction plots, and P-value tables translates "statistical significance" into "how to set parameters."

5. Residual Diagnostics: Don't Let an Outlier Skew the Entire Experiment

The analysis is not complete yet. After building the model, residual diagnostics must be performed to check three things: whether the residuals are approximately normal, whether they fluctuate with the fitted values, and whether there are any outliers.

In this case, an "insider" was identified: the defect rate in the 13th experiment was 4.1%, significantly deviating from the model's prediction. Reviewing the experiment log revealed that the flux spray nozzle was temporarily clogged that day, and the operator recorded the anomaly but did not take it seriously. Without residual diagnostics, this point would have incorrectly made "flux spray volume" a significant factor, misleading subsequent decisions. By identifying and confirming the cause through the log, and then removing the point, the model became clean. This also reminds us: experiment logs are as important as randomization sequences, and outliers are not "dirty data" but clues carrying information.

6. Confirmation Experiments and Engineering Trade-offs: Significance Does Not Mean Blind Adoption

The model suggests the optimal combination: solder temperature of 265°C, chain speed of 0.8 meters/minute, preheating temperature at the midpoint, and flux spray volume at the midpoint. The team conducted 5 days of confirmation experiments with these parameters, resulting in defect rates of 0.6%, 0.8%, 0.7%, 0.9%, and 0.7%, with an average of 0.74%, stabilizing the yield at 99.3%, achieving and exceeding the 1.0% target.

However, engineering decisions do not stop at statistics: a slow chain speed means a production capacity reduction of about 15%, which the workshop initially opposed. The finance department calculated that reducing the defect rate from 2.8% to 0.7% would save approximately 800,000 yuan annually in rework, scrap, and customer claims, far outweighing the production capacity loss. The new parameters were thus standardized and documented in the work instructions and control plan. Additionally, the center point experiment results were consistent with the linear model's predictions, indicating no significant curvature, and thus no need for a second round of optimization using response surface methodology, saving another set of experiments.

7. Retrospective: Five Key Points for DOE Result Analysis

Looking back, the truly challenging part of this improvement was not completing the 18 experiments but correctly interpreting the 18 numbers. Five key points are worth remembering:

  1. Do not use average value comparisons to replace ANOVA, as interaction effects will be averaged out.
  2. When interaction terms are significant, main effects cannot be interpreted independently.
  3. Interaction plots must be viewed alongside P-values.
  4. Residual diagnostics can identify outliers, preventing false significance.
  5. Statistical conclusions must be verified through confirmation experiments and balanced with cost and production capacity considerations.

The same experiment can yield a defect rate of 0.7% for those who know how to analyze it, while those who do not get a 2.5% defect rate and question the method. The difference in DOE often lies not in the design but in the analysis.


The first half of DOE is designing the experiment, and the second half is correctly interpreting the data; misinterpreting the data can waste even the most intricate design.

Knowledge code: 6.4.1

Version: v20260825

Author: Quality Think Tank Quality Think Tank is dedicated to providing systematic professional knowledge, methodologies, and practical tools to quality management practitioners, helping companies continuously improve their quality capabilities.