Customer Quality and Field Service Series Issue 2: Field Failure Analysis — From "Replace and Forget" to "Root Cause Closure" Management Advancement

By: QTank Published: 6/1/2026 Views: 153
Current rating: ★★★☆☆ Rate this Equivalent to 8 ratings

1. Introduction

The story doesn’t end when a product is delivered to the customer.

Every returned faulty item, every customer complaint, and every field failure report is a "health check signal" for the company’s quality system. However, in many companies, field failure handling remains at a basic level—quick to replace, but shallow in root cause analysis, leading to repeated issues and a gradual erosion of customer satisfaction through repeated complaints.

Field Failure Analysis (FFA) is the key management activity that upgrades this passive response to proactive improvement.

2. What is Field Failure Analysis? How Does It Differ from Laboratory Analysis?

Field Failure Analysis refers to the entire process of systematically collecting data, analyzing root causes, and implementing improvement loops for failures that occur in products already delivered to customers and used in real-world conditions.

It fundamentally differs from design verification or incoming quality control (IQC) conducted in a laboratory:

Dimension Laboratory Analysis Field Failure Analysis
Sample Source Test samples under controlled conditions Faulty items from actual usage environments
Stress Conditions Known, controllable Unknown, variable (temperature, vibration, operating habits, etc.)
Analysis Purpose Verify if the design meets specifications Determine why the product failed in actual use
Improvement Loop Typically limited to the design phase Can involve the entire chain from design, manufacturing, inspection, to after-sales

In short: Laboratory analysis tells you "whether the product works," while field failure analysis tells you "why it failed at the customer’s site and how to prevent the next customer from encountering the same issue."

3. Three Levels of Field Failure Analysis

Level One: Failure Data Collection

This is the most fundamental step, yet it is also the easiest to handle superficially. Many companies have records of returned items, but the data quality is poor—fault descriptions are vague, such as "broken," "not working," or "strange noise," which do not support any meaningful analysis.

Effective failure data collection should at least include:

  • Product Information: Model, batch number, production date, serial number
  • Usage Information: Installation date, operating duration, operating environment (temperature/humidity/load)
  • Failure Information: Failure mode (specific phenomena), failure time, whether it has been repaired
  • Customer Information: Customer type, industry, geographical location

It is recommended to establish a standardized Field Failure Information Registration Form for frontline after-sales personnel to fill out, avoiding subjective descriptions.

Level Two: Failure Physics Analysis

After receiving the faulty item, conduct the analysis following the principle of "first appearance, then microscopic; first non-destructive, then destructive":

  1. Visual Inspection: Take photos and observe for obvious signs such as cracks, deformation, burn marks, or corrosion.
  2. Function Reproduction: Reinstall the faulty item on the test bench to confirm if the failure mode can be reproduced.
  3. Non-Destructive Testing: X-ray, CT scanning, ultrasound—check for internal structural abnormalities.
  4. Semi-Destructive/Destructive Analysis: Cross-sectional analysis, SEM/EDS spectroscopy, metallographic analysis—determine the failure mechanism (fatigue, overload, corrosion, wear, etc.).
  5. Comparison with Similar Items: Compare with normal items from the same batch or production period.

Note that not every failure requires analysis "down to the atomic level." For low-frequency, low-impact failures, stopping at step two "function reproduction" is sufficient. The depth of analysis should be proportional to the risk level of the failure.

Level Three: Root Cause Tracing and Improvement Loop

After identifying the failure mechanism, return to the manufacturing chain and ask three questions:

Question One: Design Phase — Does the design specification cover this operating condition?

  • If the design did not consider this failure mode, update the DFMEA.
  • If the design considered it but the margin was insufficient, trigger a design change.

Question Two: Manufacturing Phase — Did the manufacturing process introduce defects?

  • Trace the production records, inspection records, and equipment parameters of the batch.
  • Check if the procedure documents clearly specify control methods for this characteristic point.
  • Verify if there were any 4M changes that were not validated.

Question Three: Inspection Phase — Can the current inspection methods detect this failure?

  • Is the sampling plan for factory inspection reasonable?
  • Is it necessary to implement full inspection or more reliable testing methods?

Only by answering these three questions can the field failure analysis truly form a closed loop.

4. Management Mechanisms for Field Failure Analysis

A good field failure management system requires the support of the following four mechanisms:

4.1 Tiered Response Mechanism

Establish a tiered response standard based on the severity and frequency of failures:

Level Criteria Response Requirement Closure Time Limit
Level A Safety-related / batch failures / major customer complaints Cross-functional team established within 24 hours 30 days
Level B Affects function but not safety Analysis initiated within 48 hours 60 days
Level C Occasional, cosmetic, or non-critical function Included in monthly statistical analysis 90 days

4.2 Failure Item Management

Create a "failure item repository" where all analyzed faulty items are retained for at least six months, along with their analysis report labels. This not only aids in cross-referencing similar failures but also serves as excellent training material for new engineers.

4.3 Failure Database

Enter data, photos, analysis conclusions, improvement measures, and verification results from each failure analysis into a database. When similar failures occur again, historical records can be directly retrieved for comparison, significantly reducing analysis time.

4.4 Monthly Failure Review Meeting

Each month, the quality department should lead a failure analysis review meeting involving design, process, after-sales, and procurement departments:

  • Review the progress of all Level A and B failure analyses from the previous month.
  • Confirm whether improvement measures have been implemented.
  • Assess if there are any abnormal changes in market failure trends.
  • Formulate meeting minutes and track pending items.

5. Common Misconceptions

Misconception One: "Only investigate returned items, not those not returned" Field failure analysis should include failures reported by customers but not returned, as well as suspicious cases identified by field service personnel. Waiting for items to be returned often misses the optimal window for collecting on-site information.

Misconception Two: "Analyzed a lot, made many changes, but the same issues keep recurring" This indicates that the improvements have not been standardized. Each root cause should be matched with specific document modifications—FMEA updates, control plan adjustments, inspection standard changes, and work instruction supplements. Any improvement not documented is not a true closed loop.

Misconception Three: "Failure analysis is the responsibility of the quality department" The root causes of field failures often lie in the design and manufacturing phases. If only the quality department is driving the process, it often lacks the necessary resources from the design and manufacturing departments, leading to incomplete actions. A cross-functional failure analysis team must be established, and the quality department should be granted the "right to halt"—the issue cannot be easily closed until the root cause is identified.

6. Conclusion

Field Failure Analysis is a crucial indicator of a company’s quality management maturity. It not only tests the analysis equipment and techniques but also the organization’s systemic capabilities—whether it can transform information from a single failure into improvement actions across the entire chain.

The difference between "replace and forget" and "root cause closure" is not in technology but in management mechanisms.

When every returned faulty item is not wasted and every failure drives improvements in design and manufacturing, the company’s quality competitiveness truly extends from the "manufacturing end" to the "usage end."

Knowledge Number: 10.2.2

Version: v20260601

Author: Quality Excellence Think Tank