Machine Design #47: Spares and Wear Parts — Choose by Criticality, Not Feeling
“Keep one of everything in stock” sounds safe, but it creates a warehouse full of uncertainty. Some items stop production immediately, some have a repair route, some wear progressively, and some are standard parts that can be sourced in hours.
A useful spare strategy answers:
What happens when this item fails, how exposed is recovery, and which support action reduces the risk at acceptable cost?
The answer belongs to the machine design, configuration baseline, maintenance task, supplier plan, and handover—not only to purchasing.
1. Where does “buy one of everything” fail?
It ignores consequence, failure probability, lead time, repairability, shelf life, commonality, and the actual configuration. It may buy cheap items that are available everywhere while missing a calibrated module with a long qualification lead time.
Inventory can also become stale: the part is obsolete, the revision is wrong, the firmware is incompatible, or no one knows which asset it fits. More boxes do not automatically mean more recoverability.
2. Do not call everything a spare
Critical spare
Failure can stop a critical function or create a safety or quality risk, and recovery cannot be achieved quickly through another route.
Insurance spare
A low-frequency but high-consequence item held to protect against a long outage or external disruption.
Wear part
An item expected to degrade through use, contamination, cycles, or environment and replaced according to condition or an evidence-based interval.
Consumable
Material consumed by the process, cleaning, lubrication, or routine service.
Repairable item
An item with a defined return, diagnosis, repair, qualification, and re-entry route.
Standard commercial part
An item with reliable availability and compatibility that may not need on-site stock if the recovery time meets the requirement.
Classification prevents a single inventory rule from being applied to incompatible risks.
3. Criticality equals consequence times recovery exposure
Consequence when the item cannot perform
Assess safety, people, equipment, quality, environment, production, data, customer commitment, and the ability to operate in a reduced mode. A small sensor can be critical if it controls a permissive or a quality release.
Recovery exposure
Assess probability, failure detectability, lead time, supplier dependency, repair time, qualification, access, calibration, shipping, customs, and the availability of a workaround or redundant path.
Use a transparent scale and record the rationale. Criticality is not a feeling and should be revisited after incidents or design changes.
4. From criticality to a support strategy
Possible actions include on-site stock, regional stock, supplier consignment, repair exchange, redundancy, a temporary workaround, condition monitoring, planned replacement, a qualified alternate, or a documented lead-time acceptance.
The strategy can differ by asset, site, and lifecycle stage. A new machine may need commissioning spares; a mature fleet may benefit from commonality and repair pooling.
5. What data is needed for stock quantity?
Use failure history, installed population, duty cycle, lead time, repair turnaround, minimum order, shelf life, service level, downtime cost, commonality, and confidence in the data. Separate demand for planned wear from random failure.
Do not pretend a precise quantity is scientific when the failure data is weak. State assumptions, review dates, and the consequence of stockout.
6. Do not replace wear parts only by calendar
Calendar replacement is useful when degradation is time-dependent and evidence supports it. It can be wasteful when wear follows cycles, load, contamination, temperature, or process material. Condition indicators, trend limits, inspection, and task evidence may produce a better interval.
If a part is replaced early, record the reason. If it fails before the interval, review the mechanism, environment, installation, and specification rather than simply shortening the calendar.
7. A repairable item needs a closed loop
Define remove, identify, package, ship, diagnose, repair, test, calibrate, quarantine, accept, and return steps. Track serial number, failure symptom, repair revision, test evidence, and warranty. A repaired unit without qualification is not automatically a spare.
8. Part identity must link to the configuration baseline
Record manufacturer, part number, revision, firmware, rating, connector, calibration, approved alternatives, compatible machine and option, drawing, BOM, and source. The installed part and the stocked part must be distinguishable.
A supplier may change an item while retaining a commercial number. Review the technical delta and update the baseline or qualification record.
9. Shelf life, storage, and preservation
Check batteries, seals, adhesives, lubricants, sensors, electronics, ESD, humidity, temperature, packaging, shock, and calibration expiry. Label receipt, lot, expiry, storage condition, and inspection status. A spare that cannot be trusted after years in a cabinet is not protection.
10. Obsolescence begins before “discontinued”
Watch last-time-buy notices, firmware support, tool availability, certifications, supplier health, component shortages, and alternate qualification. A plan can include redesign, common platform, life-buy, repair, retrofit, or a controlled migration.
Link obsolescence to the machine lifecycle and configuration baseline, not to a surprise email from a vendor.
11. Example: a hypothetical sensor module
The module controls an inspection gate, has a four-month lead time, requires calibration, and is used on three assets. Consequence is high because the line cannot release product without it. Recovery exposure is high because no alternate is qualified.
The strategy may be two calibrated modules, one in service and one protected, a repair exchange, calibration instructions, firmware compatibility, a test jig, and a migration plan. The quantity follows the risk and recovery route, not a universal “one spare” rule.
12. What belongs in the handover support package?
Include an item list, criticality and rationale, BOM and asset mapping, approved alternatives, supplier and lead time, shelf life, storage, repair route, calibration, installation and verification instructions, failure symptoms, firmware/tool dependencies, emergency contact, obsolescence status, and a baseline reference.
Maintenance should be able to identify the item, order it, install it, verify it, and update the record without reverse-engineering the project.
13. Quick review checklist
- [ ] Parts are classified by function and support route.
- [ ] Criticality records consequence and recovery exposure.
- [ ] Stock quantity assumptions and service target are visible.
- [ ] Wear intervals use evidence, condition, or a justified calendar.
- [ ] Repairable items have a closed qualification loop.
- [ ] Part identity links to configuration and asset.
- [ ] Shelf life, storage, ESD, calibration, and preservation are controlled.
- [ ] Obsolescence is monitored before discontinuation.
- [ ] Handover includes installation, verification, supplier, and emergency data.
Conclusion
Spare and wear-part strategy is a recovery design problem. Classify the part, assess consequence and exposure, select a support action, link identity to the baseline, protect storage and qualification, and keep the lifecycle visible.
The best inventory is not the largest one. It is the smallest, trusted support system that lets the owner recover the machine safely and predictably. Review the strategy after incidents, supplier changes, and lifecycle gates; criticality is a living engineering decision.
References
- ISO 55001:2014 — Asset management systems: https://www.iso.org/standard/55088.html
- ISO 10007:2017 — Guidelines for configuration management: https://www.iso.org/standard/70400.html
View all MINATA technical articles