Repeat faults are not random failures. In a warehouse automation environment, a repeat fault is a machine telling the same story twice because the maintenance system only listened to the ending. Commissioning and acceptance are the moments when that story becomes part of the permanent record, for better or worse. If the first occurrence is coded vaguely, repaired symptomatically, and documented with an incomplete picture, the second occurrence is effectively guaranteed. This article describes a practical commissioning and acceptance checklist for eliminating repeat faults before they become monthly work orders, with emphasis on inspection design, failure coding, spares strategy, condition evidence, and the decision boundaries that separate a genuine fix from a temporary reset.
Nothing in this guidance overrides site procedures, lockout requirements, OEM documentation, or competent engineering judgment. The goal is not to introduce new safety rules, but to make the data that surrounds commissioning and acceptance structured enough that repeat faults lose their hiding places.
The Usual Origin of a Repeat Fault #
Most repeat faults begin during the acceptance phase, not during operation. There is an understandable pressure to see a conveyor segment, a sorter, or a goods-to-person station produce its throughput target. When the machine trips during the acceptance test, the fastest path to completion is often to clear the alarm, reset the drive, and rerun the sequence. If it passes on the second or third attempt, the installation is accepted, the fault is closed as transient, and the site moves on.
The problem is that a transient alarm is rarely transient at the component level. An overcurrent trip may be caused by a mechanical bind that only appears under full load. A photoeye misfire may be caused by beam spread from a nearby reflective surface that appears when an operator stands in a specific position. A communication dropout may be caused by a loose termination that only changes resistance when equipment vibrates at running speed. None of these are electrical noise. They are physical conditions that were present during the acceptance test and were recorded as single-event anomalies.
To eliminate repeat faults, the commissioning process must deliberately slow down at the first unexpected event. That first event is not a nuisance. It is a clue. The acceptance team needs to decide whether the fault is a discrete one-off, a pattern waiting to repeat, or a symptom of a larger commissioning gap. Without a formal decision point, the default answer is always one-off, because that is what allows the schedule to move forward.
Commissioning as a Reliability Data Event #
Commissioning and acceptance should be treated as a reliability data event rather than a mere electrical and mechanical handover. The equipment being commissioned will spend the next ten to fifteen years generating maintenance history. The quality of that history is largely determined by how events are named and categorized during the first weeks on site.
A repeat fault becomes identifiable only when two or more events can be grouped by a common cause. That grouping is impossible if the first event is recorded as “M1 fault” and the second as “conveyor fault” and the third as “no fault found after restart.” The asset may be the same, but the failure coding makes each event look unrelated.
Failure Coding Before First Energization #
Failure coding should be designed before the equipment is ever energized. This requires the maintenance team to examine the OEM’s default alarm text and map it to their own coded taxonomy. For a warehouse conveyor system, the taxonomy should be able to express the physical object, the affected subsystem, the failure mechanism, and the operational context. A single free-text field such as “fault logged” is worse than no field, because it gives maintainers the illusion of documentation while producing no usable analysis.
At minimum, every fault record during commissioning should be able to answer four questions:
- What physical item failed or alarmed? (asset ID, subassembly, component)
- What was the observable symptom? (trip, loss of signal, jam, wrong position, slow response)
- What was the probable mechanism? (wear, misalignment, contamination, loose connection, overload, software conflict)
- What was the operational context? (empty, full load, startup, steady state, reversing, E-stop reset)
If the maintenance management system cannot capture these four fields, the commissioning checklist must include a temporary spreadsheet or logbook that does. The structure matters more than the tool. A repeat fault eliminated after commissioning is a direct result of a pattern that was visible in the data because someone coded it consistently.
Inspection Design That Finds the Second Fault, Not Just the First #
Inspection design during commissioning is often limited to checking that safety functions work and that sensors are in the right position. For repeat fault elimination, inspection must also be designed around the failure mechanisms that are most likely to be misdiagnosed. This requires a different kind of checklist, one that asks not only “does this work?” but “what would make this fail intermittently?”
The following table illustrates the relationship between a typical commissioning question, the repeat-fault symptom that appears later, and the inspection step that should have been performed.
| Typical Commissioning Check | Later Repeat-Fault Symptom | Inspection Step That Would Have Revealed It |
|---|---|---|
| Motor runs and reaches speed | Intermittent overcurrent trip under partial load | Record current draw across full range of product positions and conveyor loading states |
| Proximity sensor detects a pallet | False detection when an adjacent machine vibrates | Check sensor bracket stiffness and mounting orientation against reflective surfaces or metal edges |
| Communication link is established | Device drops off network at irregular intervals | Tap and flex the cable assembly, inspect shield termination and gland contact, then run a continuous ping test for several hours |
| Gearbox is filled with oil per nameplate | Bearing temperature climbs after two weeks of operation | Verify correct oil grade for ambient conditions and confirm filling was performed with the breather open |
| Safety gate opens and stops the machine | Random reset faults and “gate open” alarms during normal operation | Check gate alignment, hinge wear, switch actuator depth, and vibration at the gate latch point |
The common thread is that the inspection step goes beyond binary operation. It asks for a condition-based observation that can be compared to a baseline. When the first fault later appears, the maintenance team can consult that baseline and decide immediately whether the fault is a variation of a known condition or a new phenomenon.
Condition Evidence and the First-500-Hour Baseline #
Repeat faults are often missed because there is no baseline evidence against which to compare a later observation. During commissioning, the team has a unique opportunity to capture condition data that will never be easy to capture again, because the system is new, open, and accessible.
The first-500-hour baseline should include at least the following types of evidence:
- Vibration measurements at drive end and non-drive end of motors, on pump casings, and on gearbox input and output bearings
- Temperature readings on motor housings, gearbox surfaces, and variable frequency drive heatsinks under normal load
- Current recordings on each phase of every drive under steady-state load and during acceleration
- Photographs of alignment, belt condition, chain sag, and sensor positions before covers are fitted
- Timing logs that show scan cycles, communication latency, and sequence completion times
- If the system can interface with higher-level controls, the exact text of every alarm and event message as it appears on the HMIs and in the historian
This evidence creates a reference point. A heat signature that is 10°C higher than the baseline during normal operation is an early indicator of a repeat fault that has not yet fully expressed itself. Without the baseline, the same reading may be dismissed as “probably normal for the running conditions today.” That dismissal is how the second or third occurrence of a mechanical problem becomes a pattern that takes three months to diagnose.
For condition evidence to be useful, it must be stored in a way that is accessible to the maintenance team that will actually work on the machine. A paper copy in the OEM manual is less useful than a digital folder named by asset ID with the date of collection. The expectation is not that every conveyor will have a permanent vibration monitoring system. The expectation is that the commissioning team will take a defined set of measurements with portable instruments and leave them in a known location.
Common Interpretation Errors in the Acceptance Window #
There are several interpretation errors that recur across sites when commissioning data is reviewed. Recognizing them helps the maintenance and controls teams avoid building a repeat fault into the system design.
Correlating speed with health. A machine that runs fast is not necessarily healthy. During acceptance, a conveyor may run product faster than its rated speed because the frequency inverter was set to 60 Hz on a 50 Hz motor, or because the PLC accelerates the drive more aggressively than the mechanical design allows. The machine passes the throughput test, but every cycle produces additional bearing wear and chain tension. The resulting repeat fault is always categorized as mechanical wear rather than commissioning configuration.
Interpreting a successful retry as a resolution. When a machine faults and then runs correctly after a reset, it is natural to conclude that the fault was transient. But a successful retry only proves that the condition that caused the fault is intermittent. It does not prove that the root cause is gone. The correct interpretation is that a condition exists that is sensitive to some variable, possibly time, temperature, load, or position. The retry does not eliminate the variable, it only demonstrates that the machine passed through a favorable state.
Replacing a sensor without checking its mounting. This is the single most common repeat fault generator in warehouse automation. A photoeye fails, a spare is installed, the copy runs correctly, and the fault is closed. The same photoeye fails three weeks later. The root cause is vibration at the mounting bracket, or contamination accumulating on the lens due to airflow, or a reflective surface moving into the beam path. The sensor replacement is not a repair. It is a temporary reset that removes a component while leaving the environmental condition in place.
Treating PLC logic as infallible. During acceptance, faults that originate in control logic can look like hardware faults. For example, a palletizer might stop because the PLC decided a clamp cylinder did not reach its home position within the allowed time. The physical cylinder is fine, but the previous cycle did not fully retract because a mechanical stop was not adjusted. The PLC then times out and reports a fault on the cylinder sensor. If the sensor is replaced and the machine is reset, the underlying mechanical adjustment error remains and will repeat at a predictable cycle count.
These interpretation errors are not caused by poor technicians. They are caused by a lack of decision structure in the acceptance process. When the process does not require the team to define a root cause before resetting, the team will naturally reset and move on.
Spares Strategy Tied to Failure Codes #
Repeat fault elimination has a direct relationship with spares management. The maintenance team cannot correct a recurring failure mode if the spare part is not available or if the wrong spare is stocked. Commissioning is the time to compare the asset list against the recommended spares in the OEM documentation and to adjust for the failure codes that emerge during the first few weeks of operation.
For warehouse automation, the highest-value spares during the first year are not always the most expensive components. They tend to be the components that are subject to environmental conditions, adjustment, and wear, such as photoeyes, limit switches, prox sensors, v-belts, bearings, and drive belt segments. However, the presence of a spare is only useful if the maintenance team knows what failure mode the spare is intended to cover.
It is useful to create a simple spares decision matrix during commissioning, based on the failure codes observed:
- If the failure code indicates contamination, identify the source and determine whether a sealed sensor, an air purge, or a repositioned guard is more effective than a spare part.
- If the failure code indicates misalignment, a replacement part will fail again unless the alignment procedure is written and verified.
- If the failure code indicates overheating, the spare part must be the correct thermal rating for the installation, not the rating that is easiest to source.
- If the failure code indicates no fault found, the spare part is irrelevant until the root cause is found. Do not add a spare to stock for a failure mode that has not been defined.
This approach prevents the classic scenario of a high-cost emergency spare that is purchased after a system stoppage but never used, because the true failure mechanism is a cheaper but less obvious component. The acceptance checklist should include a review of spares against the actual failure codes recorded during commissioning, not only against the OEM recommended spares list.
Decision Boundaries: Intervene, Repair, Redesign, or Observe #
Not every anomaly discovered during commissioning requires immediate correction. The checklist must define decision boundaries so that the team does not oscillate between overcorrecting and undercorrecting.
Intervene when the observed condition presents an immediate risk to personnel, product, or the asset itself. Examples include a guard that vibrates loose, an overheating gearbox, or a sensor position that causes intermittent safety gate signals. In those cases, corrective action is mandatory before the machine is released to production.
Repair when the condition is a clear deviation from the baseline and a known fix exists. For example, a drive shaft coupling that was installed without the correct gap will cause repeated vibration. The repair is to set the gap per OEM procedure. The condition should be fixed at commissioning and the baseline should be updated.
Redesign when the condition causes a repeat fault and the part itself is unlikely to survive in the environment. This is common with sensors placed where fork trucks or maintenance personnel regularly bump them, or with cable carriers that flex beyond their designed radius. A redesign may be as simple as installing a protective bracket or as involved as relocating the sensor. The decision boundary is met when a second occurrence appears on the same component within a short period.
Observe when the condition is present but cannot be correlated with a likely failure mechanism. In this case, the checklist should require a documented observation plan with a specific timeframe, a signal that will trigger action, and the person responsible for reviewing the data. Uncontrolled observation, by contrast, is a polite word for doing nothing.
The decision boundary that must be respected is the one that separates observation from delay. If the team chooses to observe, it must define what would change their decision. For example, if a bearing temperature runs 8°C above baseline, the team might decide to observe for 100 operating hours and intervene if the temperature rises another 5°C. That is a measurable boundary. Without a boundary, observation becomes procrastination.
The Acceptance Checklist Structure #
The acceptance checklist for eliminating repeat faults is not a single list that is completed at the end of installation. It is a staged set of checks that begins before energization, continues through the first-500-hour baseline, and is only closed after the machine has operated through at least one full load cycle without needing a reset.
A workable structure has four stages:
Stage 1: Pre-energization data preparation. Confirm that the failure coding taxonomy exists, the asset database is populated, and the condition baseline locations are defined. Verify that the maintenance team has read access to the HMI and controls historian and that alarms can be exported in a format that can be reviewed.
Stage 2: First-run functional verification. Run each machine at low speed and then at rated speed. Record current, temperature, and vibration at each stage. Log every alarm, including alarms that appear, clear, and reappear. Do not clear a fault to make a test pass without writing a note explaining the fault and the likely mechanism.
Stage 3: Loaded operation and simulated production. Introduce real product or equivalent load at the expected cycle rate. Observe the machine for at least one complete shift, including startup from cold and restart after a planned stop. Note any condition that changes when the machine is warm, when the product queue is full, or when multiple machines run simultaneously.
Stage 4: First-500-hour review. At approximately 500 operating hours, repeat the condition measurements from Stage 2 and compare them to the baseline. Review all fault records and group them by the coded mechanism. Remove any spare part from stock that was purchased for a fault that never repeated, and add any part that would have reduced downtime for a fault that did repeat.
The stage structure is only effective if the data is actually reviewed. A completed checklist that is filed without analysis is indistinguishable from no checklist. The final step of the commissioning process should be a formal review meeting where the failure codes, condition evidence, and spares adjustments are presented to the operations and maintenance leadership.
Key Takeaways #
- Treat every first fault during commissioning as a clue, not a nuisance. The decision to reset without root cause analysis is the most reliable predictor of a repeat fault.
- Design the failure coding taxonomy before energization so that the first event and the tenth event can be grouped by asset, mechanism, and context.
- Use condition evidence from commissioning as the baseline for all later comparisons. Vibration, temperature, current, and alignment data taken on day one turn the second fault into a measurable deviation rather than a mystery.
- Build inspection checks that challenge the machine, such as flexing cables, checking bracket stiffness, and monitoring current under varying load, instead of only checking for binary operation.
- Do not interpret a successful retry as a resolution of the root cause. A retry proves the machine is intermittent, not that it is fixed.
- Align the sp
Related Pearl Gateway Guides #