Machine Design #40: Mechanism FMEA — From Failure Mode to Design Action
A mechanism FMEA should not end as a spreadsheet full of severity, occurrence, and detection scores. Its value is the chain from a real function, through a credible failure, to a design action that changes the machine or produces evidence that the risk is controlled.
If the team starts by listing purchased components, the analysis often misses interfaces, operating context, stored energy, human access, and failures that occur only during recovery or maintenance. A good DFMEA is a reasoning tool for design review. It shows what the mechanism must do, how it can fail, what the consequence is, who owns the action, and how the action will be verified.
This article is a conceptual engineering guide. It does not replace a formal risk assessment, applicable safety standards, customer requirements, or validation by competent engineers.
1. Start before the design is frozen
Open the FMEA when functions and interfaces are still changeable. A review after drawings, tooling, and software are frozen turns the document into an archive of problems that the team can no longer afford to fix.
At concept, use assumptions and mark them visibly. As the design matures, replace each assumption with an analysis, calculation, test, supplier document, or approved decision. The FMEA should evolve with the design baseline, not be written once at the end.
2. Choose scope by function, not by parts list
Define the boundary around a mechanism and its interfaces: locating, clamping, transfer, lifting, pressing, cutting, or guarding. Include the energy sources, workpiece, sensors, utilities, operator interaction, and neighboring stations that influence the function.
A parts list tells you what was purchased. A functional scope tells you what can go wrong. Keep the boundary narrow enough for a useful review but wide enough to include incoming forces, feedback, and recovery.
3. Write the function before searching for failure
State what the mechanism must do, under which conditions, with what limits and evidence. For example: “Locate model B within 0.10 mm, resist the defined process force, release without damaging the part, and provide a confirmed state before the next operation.”
Avoid writing “cylinder works” or “sensor ON.” Those are implementation fragments, not functions. A function gives the team a reference for judging partial performance and failure effects.
4. Separate failure mode, effect, and cause
- Failure mode: what fails locally—does not clamp, clamps in the wrong position, leaks, cracks, loses alignment, or gives a stale signal.
- Effect: what the failure causes at the mechanism, machine, product, operator, or downstream level.
- Cause: why it occurs—wear, overload, contamination, wrong tolerance, incorrect assembly, software state, utility loss, or interface mismatch.
Do not write a cause as a failure mode or hide several different failures in “mechanism error.” Different causes need different prevention and detection controls.
5. Review effects at several levels
Describe the local effect first, then propagate it. A loose locating pin may create an incorrect datum, a dimensional defect, a downstream jam, a customer escape, or a person entering a hazardous zone to clear it. The same local failure can have different effects in automatic production and manual recovery.
Separate immediate effect, next-operation effect, system effect, and human-safety effect. This prevents a high-consequence failure from receiving a low score merely because the local component still moves.
6. Look for failures at interfaces
Interfaces commonly hide failure: bracket-to-frame stiffness, cylinder-to-clamp alignment, sensor-to-target distance, robot-to-fixture handoff, PLC-to-drive handshake, and upstream-to-downstream timing.
Ask what happens when one side is within its own specification but the combination is not. A connector can be correctly wired yet mapped to the wrong model. A conveyor can meet its speed requirement while arriving before the receiving clamp is ready. Interface requirements need an owner on both sides.
7. Prevention control is different from detection control
Prevention reduces the probability of the failure: correct datum design, adequate stiffness, poka-yoke, controlled tolerance, contamination protection, overload limitation, and proper assembly.
Detection identifies a failure after it has occurred or while it develops: position feedback, force signature, pressure switch, vision check, cycle-time trend, or proof test. Detection is not automatically a substitute for prevention. A sensor that stops a bad part is valuable, but it does not make a weak mechanical design acceptable without considering escape and response.
8. Do not let a total score hide a critical failure
Risk-priority numbers can support discussion, but multiplication can hide a failure with high severity and low occurrence. Review severity, occurrence, detection, exposure, and safety significance separately where required.
Define escalation rules: any credible injury mechanism, uncontrolled stored energy, single-point loss of containment, or inability to detect an unsafe state should receive engineering attention regardless of a favorable total score. Explain why a risk is accepted and what evidence supports the decision.
9. An action must change design or evidence
“Review,” “be careful,” and “monitor” are not complete actions. A useful action names the owner, due date, affected item, decision, and verification evidence.
Examples: increase guide spacing and verify stiffness by calculation; add a mechanical catch and proof-test the holding load; change sensor placement to measure seating and challenge it with debris; add a pressure diagnostic and record the response time; update the recovery sequence and test power loss at maximum load.
Close an action only when the design or controlled evidence changes. Attach the drawing revision, calculation, test record, supplier document, or approved rationale.
10. Example: FMEA for a workpiece stop assembly
Function: stop a pallet at a defined datum, withstand the impact and process force, confirm the stop state, and retract without leaving a collision hazard.
| Failure mode | Effect | Credible cause | Prevention | Detection | Design action |
|---|
| Stop pin bends | Datum shifts; part defect | Impact energy above rating | Size for worst-case energy; add guide | Position check and dimensional audit | Recalculate energy and validate at boundary |
| Stop does not extend | Pallet passes; collision downstream | Air loss or valve failure | Load-holding architecture; protected routing | Extend sensor and timeout | Add pressure diagnostic and controlled recovery |
| Stop remains extended | Transfer collision | Return spring broken; stale output | Mechanical clearance and fail state | Retract sensor; handshake timeout | Add independent clearance check |
| Sensor reports false clear | Cycle advances with hazard | Target loose or connector fault | Rigid target and keyed wiring | Plausibility and challenge test | Add diagnostic state and replacement check |
The table is useful only if each action reaches a drawing, parameter, test, or approved decision. A score without an action is not risk reduction.
11. Link FMEA to other deliverables
Connect the FMEA to requirements, risk assessment, interface control documents, drawings, calculations, software states, alarm definitions, test cases, spare-parts lists, and maintenance instructions. Use stable IDs so an action can be traced from failure mode to evidence.
When a design changes, identify which FMEA rows, requirements, tests, and training documents are affected. A baseline that cannot show these links creates false confidence during handover.
12. When must the FMEA be updated?
Update it when the function, load, material, supplier, tolerance, software sequence, sensor, safety measure, environment, maintenance interval, or recovery method changes. Also update it after a recurring field failure, near miss, rejected part, abnormal test, or important assumption is disproved.
Do not wait for the next formal review. A change that affects a failure path should trigger impact review before release.
13. Design-review checklist
- [ ] Scope is defined by function and includes energy, interfaces, people, and recovery.
- [ ] Each function states operating context, limits, and required evidence.
- [ ] Failure mode, effect, and cause are separate and specific.
- [ ] Effects are considered locally, downstream, at system level, and for human safety.
- [ ] Interface failures and common-cause conditions are reviewed.
- [ ] Prevention and detection controls are distinguished.
- [ ] High-severity or single-point failures are escalated beyond a total score.
- [ ] Every action has an owner, due date, affected design item, and verification route.
- [ ] Closed actions reference a drawing, calculation, test, or approved rationale.
- [ ] FMEA rows link to requirements, risk assessment, alarms, and test evidence.
- [ ] Changes, field failures, and disproved assumptions trigger updates.
Conclusion
Mechanism FMEA is valuable when it changes engineering decisions. Start from function, follow energy and interfaces, separate failure mode from effect and cause, and judge consequences at the level where people, product, and the machine are actually affected.
Use scores to focus discussion, not to hide critical failures. Convert every important finding into a design action or controlled evidence with an owner and verification route. Then connect the FMEA to the requirements, drawings, software, alarms, tests, and maintenance documents that keep the decision alive after handover.
Public references
View all MINATA technical articles