DFMEA Missed Judgment Case Analysis —— Improving Design Risk Identification from a Batch Failure

By: QTank Published: 8/7/2026 Views: 51
Current rating: ★★★☆☆ Rate this Equivalent to 8 ratings

Abstract: No matter how beautifully a DFMEA is done, if key failure modes are missed, all preventive measures are just empty talk. This article follows a real improvement project at an automotive electronics component company, thoroughly reviewing a batch failure caused by a missed DFMEA judgment: from the sudden surge in market complaints, root cause analysis, to the systematic review of the DFMEA, identifying the missed judgment, and finally rebuilding the risk identification process and implementing preventive measures. Beyond the case, this article distills four typical modes of DFMEA missed judgments and a five-step improvement loop, helping you to verify whether your DFMEA is a "risk map" or merely a "form-filling exercise."


1. A Batch Failure Exposes the Illusion of "Zero High Risk"

A certain automotive electronics component company (hereinafter referred to as Company A) specializes in body control modules, supplying multiple vehicle manufacturers. In the third quarter of 2025, the after-sales market suddenly saw a concentrated surge in complaints: a batch of products experienced poor contact at the power terminal, leading to intermittent functional interruptions in the vehicles, with the most severe case causing a production line shutdown at a customer's facility. In just three weeks, over two thousand units were returned, resulting in direct losses and production line shutdown claims totaling nearly 80 million yuan.

The first reaction of the quality department was to investigate the manufacturing process: welding parameters, terminal crimping force, incoming material batches... But after a thorough investigation, all process data fell within control limits, and no abnormalities were found in the manufacturing process. At this point, someone raised a question: "How was the risk of this terminal assessed in the initial DFMEA?"

Upon reviewing the DFMEA document, everyone fell silent: in this fully reviewed and signed DFMEA, the failure mode of poor contact at the power terminal did exist, but the risk priority number (RPN) was only 36 points—S (severity) was rated 6, O (occurrence) was rated 2, and D (detection) was rated 3. According to the company's rule at the time, "RPN greater than 80 requires corrective action," this failure mode was deemed "low risk, acceptable," and no additional measures were taken.

A failure mode that caused a production line shutdown was rated only 36 points in the DFMEA. This is not a matter of scoring, but a systemic missed judgment in the risk identification chain at some point.

2. Root Cause Analysis: Discrepancy Between the Physical Mechanism of Failure and DFMEA Records

The improvement team did not rush to change the score but first thoroughly understood the physical mechanism of the failure using traditional quality tools.

The first step was fault tree analysis (FTA). The team took "poor contact at the power terminal" as the top event and expanded it layer by layer: possible causes of poor contact included terminal deformation, insufficient insertion force, wear of the plating, foreign object intrusion, and vibration-induced loosening. Through the dissection and analysis of the returned units, the primary cause was ultimately identified as insufficient terminal retention force. Further electron microscopy and metallographic analysis revealed that the terminal spring experienced stress relaxation under high-temperature and vibration conditions, leading to a continuous decrease in insertion force over time, eventually falling below the critical threshold.

The second step was to verify the discrepancy between the failure mechanism and the DFMEA records. After a reverse DFMEA review, three key contradictions were found:

First, the failure mode was incorrectly identified. The original DFMEA described the failure mode as "increased terminal contact resistance," while the actual physical failure was "terminal retention force decay leading to loose insertion." Although both ultimately manifest as poor contact, they have different mechanisms, influencing factors, and preventive measures—misidentifying the failure mode is like aiming at the wrong target.

Second, the triggering conditions were not identified. The position of the terminal in the vehicle was subject to long-term high-temperature and vibration conditions, but the DFMEA analysis referenced static plug-pull durability data at room temperature, failing to consider the combined conditions of "high temperature + vibration + long-term." The O value of 2 points was based on laboratory data at room temperature—actual conditions were much more severe.

Third, the detection measures were overestimated. The D value of 3 points in the DFMEA was based on "full inspection of plug-pull force at the factory." However, the full inspection used new terminals, which could not detect time-dependent failures such as "retention force decay over time." The mismatch between detection measures and failure mechanisms led to an inaccurate detection score.

A low-risk item with 36 points was actually a high-risk item that combined "misidentified failure mode, missed triggering conditions, and mismatched detection measures." The cumulative effect of these three discrepancies completely nullified the DFMEA's early warning function.

3. Reverse Comparison: Backtracking Each Step of the DFMEA from the Failure Result

After the root cause was clear, the improvement team did something many companies overlook—systematic backtracking. Using the "verification" logic between D4 (root cause analysis) and D5 (permanent corrective action) in the 8D method, they treated this failure as a "real-world drill" and systematically backtracked each step of the DFMEA to pinpoint where the missed judgment occurred.

The backtracking followed the structure of the AIAG-VDA seven-step method:

Structure Analysis Phase: The DFMEA structure tree categorized "power terminal" under "connector assembly," but did not further decompose it to the "spring" sub-component. The stress relaxation characteristic of the spring was precisely the physical root of this failure. Insufficient granularity in structure analysis led to all subsequent analyses remaining at the "terminal" level, failing to see the "spring."

Function Analysis Phase: The function description was "transmit power signals," but lacked time-based functional requirements such as "maintaining reliable contact throughout the vehicle's lifecycle." The coarser the function description, the easier it is to miss failure modes.

Failure Analysis Phase: The failure chain only went as far as "increased contact resistance" and did not further investigate "why contact resistance increases." The failure chain was broken, and the root cause (stress relaxation) never entered the DFMEA's scope.

Risk Analysis Phase: S, O, and D scores were all based on "data available at the time," without referencing real-world operating conditions and time-dependent data. The scoring became a "consensus in the meeting room" rather than a "data-supported estimate."

Optimization Phase: Because the RPN was below the threshold for corrective action, the optimization phase skipped this failure mode—no measures, no verification, no responsible person. The DFMEA loop was broken here.

The conclusion from the backtracking was alarming: the missed judgment was not a mistake in one phase but a "total collapse of the chain" from structure analysis to optimization. Each phase was slightly off, and when combined, they resulted in a batch failure.

4. Five-Step Improvement Loop: Transforming DFMEA from a "Document" to a "Risk Map"

Company A did not stop at "fixing this one failure" but established a replicable DFMEA improvement loop with five steps:

Step 1: Rewrite Failure Modes. Rewrite all critical failure modes according to the "function loss + mechanism description" standard. For example, change "increased contact resistance" to "terminal retention force decay under high-temperature and vibration conditions, leading to increased contact resistance and eventual functional interruption." The standard requires that failure modes must describe "what broke and how it broke," not just general phenomena.

Step 2: Mandatory Reference to Operating Conditions. Establish a product operating conditions database and use parameters such as temperature range, vibration spectrum, power-on time, and environmental media as mandatory inputs for DFMEA analysis. For any time-dependent failures (stress relaxation, fatigue, wear, aging), the corresponding durability data under the specific operating conditions must be referenced, and room temperature static data cannot be used as a substitute.

Step 3: Integrity Check of Failure Chains. Each failure mode must be questioned at least two levels deep to "why it happens," down to the physical/chemical mechanism level. The team borrowed the FTA approach, making "failure mode → failure cause → mechanism" a required field. Failure modes without a clear mechanism are considered incomplete. This step is the most time-consuming but also the most revealing—when Company A reviewed its existing DFMEA, nearly 40% of the failure modes could not answer the second-level question, indicating that much of the past analysis was superficial.

Step 4: Verification of Detection Measures. Each detection measure must answer "can this measure truly detect this failure?" The criterion is that the physical principle of the detection method must match the failure mechanism. Time-dependent failures must be verified using aged samples, not just assuming that "full inspection is effective."

Step 5: Regular Review of Low-Risk Items. Abolish the practice of "never reviewing items with RPN below the threshold" and change it to "quarterly rolling review of low-risk items." Review triggers include market complaints, design changes, operating condition changes, and the introduction of new materials and processes. RPN is no longer a one-time ticket but an input for dynamic risk management—risks change over time and with operating conditions, and passing one review does not guarantee eternal safety.

5. Implementation of Improvements: From Case Rectification to Systemic Capability

After establishing the five-step loop, Company A took three concrete actions:

First, comprehensive review of existing DFMEAs. According to the new standard, they reviewed the DFMEAs of all more than twenty product platforms in production, and found seven similar "low-score high-risk" items, two of which were deemed to require immediate corrective action. The cost of the review was two weeks of concentrated evaluation, but compared to a production line shutdown, this cost was negligible.

Second, verification measures and preventive actions. For this failure, they added anti-stress relaxation process treatment to the terminal spring material and established "sampling inspection of plug-pull force after high-temperature aging" as a new detection method. They also included this failure mode in the control plan and inspection specifications. Three months of tracking data showed that the retention force decay rate of the improved product in accelerated aging tests decreased by more than 60%.

Third, case-based training materials. The entire process of this failure—from the surge in complaints to the DFMEA backtracking, from the five-step loop to the verification data—was compiled into an internal case and included in the mandatory training for design engineers and quality engineers. The core of the training is one simple sentence: each DFMEA score must answer "where the data comes from, what the mechanism is, and whether the measures have been verified."

Six months later, Company A underwent another annual customer audit. The auditor, while reviewing the DFMEA, asked a question: "How do you ensure this won't happen again?" The improvement team's response was: each high-risk item in the DFMEA is supported by operating condition data, has a fully expanded failure chain, has verified detection measures, and has a record of rolling reviews—four conditions for missed judgments, and we have at least addressed three of them. The path from 36 points to a corrective action loop is precisely the complete journey of DFMEA from a "document" back to a "risk map."

6. Final Thoughts

Looking back at this improvement, the most important takeaway is not how many points the RPN was changed to, but a simple truth: the value of DFMEA lies not in how full the table is, but in whether the risk identification chain is broken. A little coarser in structure analysis, a little more general in function description, a little shorter in failure chain, a little older in data, and a little more presumptuous in detection—each "little" alone is not fatal, but they can accumulate along the chain and eventually erupt as a batch failure.

The true significance of the improvement loop is to transform "post-failure astonishment" into "pre-failure questioning": is the failure mode correctly identified? What is the mechanism? Is there data support? Is the detection effective? Will low-risk items change? By asking these four questions, the DFMEA can transform from a document for audit compliance back into what it should be—a risk map for the design team.


DFMEA never misses a specific failure mode; it misses every "almost" in the risk identification chain from structure to optimization.

Knowledge code: 8.2.1

Version: v20260807

Author: Quality Think Tank Quality Think Tank is dedicated to providing systematic professional knowledge, methodologies, and practical tools for quality management practitioners, helping companies continuously enhance their quality capabilities.