When a warehouse system is handed over for commissioning, every fault found is an invitation to ask a deeper question. A fault is not only a machine behaviour that deviates from expectation; it is also a record of how the installation, configuration, and operational environment have interacted under stress. A safe fault investigation is one that treats the recovery of the system as a secondary objective and the protection of people, evidence, and process integrity as the primary one. This article describes a disciplined approach to fault investigation during the commissioning and acceptance phase, with particular attention to access control, lockout discipline, evidence collection, and the boundaries of engineering judgement.
Purpose and Scope of Acceptance-Based Fault Investigation #
Commissioning is the period when a warehouse system passes from construction into productive operation. Acceptance is the moment when the operator confirms that the system is capable of performing its designated function within agreed limits. A fault found during this phase carries special significance because it may indicate not simply a failed component but a wrong assumption made during design, a deviation during installation, or an incomplete understanding of how the system behaves under load.
The purpose of a fault investigation during acceptance is not to assign blame. It is to build a reliable record of whether the system can run safely and as intended. Investigators must be willing to slow down, protect the scene, and consider that the fault might be the only visible clue to a hidden interaction between hardware, software, and human action.
Because the commissioning environment is often noisy, time-pressured, and populated by several trades, the investigation process must be governed by a clear scope. That scope should be defined before the interrogation begins. Questions such as whether the system is in a safe state, whether power can be isolated without losing diagnostic data, and who has authority to authorise access must all be answered in advance. Site procedures, lockout requirements, and OEM documentation are the controlling references in these situations, and the guidance offered here is intended as a supplement to them, never as a replacement.
Operating Context: Where Faults Hide in a Warehouse System #
A typical warehouse system is a network of interacting subsystems. Conveyors, sortation units, lifts, transfer cars, palletisers, and automated storage machines rely on a common foundation of safety interlocks, guard doors, emergency stop circuits, sensors, and programmable logic controllers. A fault in one subsystem frequently appears as a symptom in another.
Consider a photo-eye sensor on a conveyor that is misaligned by a few degrees. The sensor may report a clear path when a case is present, or it may report a blockage when there is none. In the first case, the downstream merge conveyor may run empty and create a timing error. In the second, the conveyor may stop and request operator intervention. Both conditions appear to be conveyor faults, but the root cause is the sensor mounting, not the conveyor drive.
In a similar way, a short circuit in a limit switch cable can feed back into the safety relay circuit and cause an entire zone to drop out. The visible symptom is a guard door that appears closed but will not reset. The investigator must be careful not to assume the guard-door switch itself is faulty simply because it is the most proximate component.
The interaction between mechanical and electrical systems is particularly important during acceptance. A roller that is slightly out of parallel may cause a case to drift, which then triggers a tracking sensor. The fault is reported as a sensor failure, but the real issue is mechanical. These cross-domain interactions are why commissioning investigations require thinking beyond the immediate fault code.
Before Touching the System: Evidence Collection and Boundary Setting #
The first and most critical rule of safe fault investigation is that nothing is touched until the system is brought to a controlled state. This means following the site procedure for controlled shutdown, locking out energy sources, and confirming that the system is isolated from pneumatic, hydraulic, and electrical power. Each organisation has its own authorised lockout steps, and those steps always take precedence over any generic guidance.
Once the system is locked out, the investigator must shift from action to observation. The goal is to collect enough evidence to reconstruct the sequence of events leading up to the fault. The following items should be documented before any adjustment, repair, or even measurement is attempted:
- Photographs of the entire affected zone, not just the failed component, including clear shots of any equipment that appears to be misaligned, displaced, or marked with impact scuffs.
- The exact text of fault messages, error codes, and operator panel prompts, recorded verbatim without interpretation.
- Time stamps and sequence logs from the PLC or SCADA system, if the system remains energised for data retrieval under controlled conditions.
- A written record of who was present, what they saw, and what they did during the last few minutes before the fault occurred.
- Any environmental observations, such as unusual temperature, humidity, or the presence of packaging debris, dust, or water near the affected area.
- A sketch of the conveyor or machine layout showing the physical relationship between the reported fault point and adjacent components, including guard doors, interlocks, and emergency stops.
Boundary setting is a separate discipline. The investigator must define which parts of the system are inside the scope of the investigation and which are outside. This boundary is not a physical line only; it also includes a time boundary. Evidence from before the last maintenance intervention may be relevant, but it must be treated with care because the system may have been changed since then.
Access control is a further boundary. Only personnel with the correct competence, training, and authorisation should enter the area of investigation. That may sound obvious, but in a busy commissioning environment, well-meaning operators and electricians often approach a fault with the intention of helping. This can destroy evidence and create additional risk. The overarching principle is that investigation is a controlled task, not a casual activity.
Observable Symptoms and Their Probable Component Interactions #
Every symptom carries meaning, but the meaning is rarely visible at the surface. The investigator learns to read a symptom as the beginning of a chain of reasoning rather than as a conclusion. For example, a conveyor zone that intermittently fails to start may be caused by a temperature-sensitive overload relay, a poor connection in an emergency stop chain, or a software interlock that is being flagged by a neighbouring zone. The symptom is the same, but the investigation path is very different in each case.
The following symptom patterns are common during warehouse system acceptance and are worth studying carefully because they reveal how different subsystems interact.
Intermittent faults are among the most dangerous to interpret. Because they appear only under certain conditions, such as particular load levels, particular carcase sizes, or particular environmental temperatures, they tempt the investigator to replace a component and then wait for the fault to return. This is often wasteful and sometimes unsafe. A better approach is to identify the condition that is present every time the fault appears and every time it is absent. This condition is usually the investigative key.
Multiple faults reported at the same time are often caused by a single shared factor. If several conveyor zones drop out simultaneously, the common element may be a single safety relay, a shared power supply, or a network cable segment. If several sensors report contradictory states, the common element may be a damaged harness or a grounding problem. Chasing each fault independently will lead to confusion; seeking the shared factor is more productive.
Faults that clear soon after being reset are another special category. These may indicate that the system is operating at the edge of its design window. For example, a lift that occasionally times out because a load is near the maximum permissible weight, or a sortation chute that jams because the divert actuator is marginally too slow, will often reset cleanly. The investigator should never accept a clean reset as proof that the underlying problem has been resolved.
Common Interpretation Errors in the Acceptance Phase #
Even experienced engineers make predictable mistakes when interpreting fault evidence during commissioning. Recognising these errors in oneself and in the wider team is a significant part of an independent investigation. The first and most common error is the assumption that the most recently changed component is the cause of the fault. In commissioning, many components are installed, adjusted, and reconfigured within a short period. A fault that appears after a specific adjustment is often attributed to that adjustment, when in fact the fault was present earlier and simply not noticed.
A second error is the fixed-state assumption, where the investigator looks at a component in its current state and assumes it has always been in that state. A misaligned sensor may have been moved during an earlier maintenance intervention, or a guard door may have been adjusted during commissioning. The current position can be misleading if the history of the component is unknown.
A third error is the symmetry error. This occurs when an investigator assumes that if component A fails and component B is identical to A, and component B appears to be working, then the fault is purely in A. This ignores the possibility that the two components are subjected to different loads or different environmental conditions. Identical components are often not identically stressed.
A fourth error is the reset-and-resume error, where an operator clears a fault to restore production without first recording the fault context. This is not a matter of carelessness; it is often a matter of pressure. However, when a fault is cleared without documentation, the most valuable evidence is permanently destroyed. A culture that values evidence over speed is essential during acceptance.
Finally, there is the complexity error, where the investigator assumes the cause of the fault must be as complex as the system itself. In practice, many faults have a simple root cause, such as a loose fastener, an incorrectly routed cable, or a sensor that has been set with the wrong sensing distance. The investigation should always begin with the most accessible and most likely causes, but without narrowing the field prematurely.
Practical Diagnostic Table: Symptom, Evidence, and Likely Interaction Points #
The following table provides a practical reference for commissioning investigators. It is not intended to replace the OEM fault chart, but it may help in the early triage phase before those charts are consulted. The investigator should use the table as a memory aid and not as a definitive guide to repair.
| Observable Symptom | Possible Component Interactions | Evidence to Collect | Typical Interpretation Pitfall |
|---|---|---|---|
| Conveyor zone will not start after reset | Guard door switch, emergency stop relay, PLC output module, contactor coil, motor thermal overload | Reset procedure used, fault code history, state of all safety relays, voltage at contactor coil | Assuming the contactor is faulty because the motor does not energise |
| Intermittent sensor loss in one zone | Photo-eye alignment, cable harness, connector, PLC input module, interference from a nearby VFD or motor starter | PLC force table or input status, wire routing, sensor mounting torque, environmental temperature | Replacing the sensor when the issue is a damaged connector pin or cable route |
| Multiple zones drop out simultaneously | Shared safety relay, common 24 VDC power supply, Ethernet/IP network trunk, central emergency stop panel | Time-synchronised logs across zones, power supply status, network diagnostic counters | Testing each zone individually and concluding that all have separate faults |
| Sortation chute jam with case present but sensor clear | Divert actuator position, pneumatic pressure, photo-eye threshold setting, case dimensions, conveyor speed | Case size and weight, actuator travel time, air pressure at the regulated value, sensor output with a known target | Adjusting the sensor before checking the actuator response time |
| Machine announces false guard door open | Guard door switch contact, actuating cam, door hinge wear, wiring short to ground, safety relay input | Switch state with door closed, door sag measurement, insulation resistance of cable, PLC input state | Replacing the switch without checking that the door itself has moved |
| Lift or palletiser stall at a high position | Encoder, load cell, frequency drive, mechanical brake, cable chain, PLC positioning logic | Encoder position at stall, drive current and frequency, brake engagement, load at time of stall | Assuming a drive fault when the load has shifted during the cycle |
When using this table, the investigator should fill in the evidence column before deciding which of the interaction points is most likely. A clear evidence record prevents the discussion from drifting into speculation.
Maintenance Implications of Acceptance Findings #
Faults found during acceptance are not only problems to be solved; they are also early indicators of how the system will behave once it enters service. A fault that is caused by a misadjusted sensor or a poorly routed cable will certainly recur unless the underlying condition is corrected. If the investigator addresses only the immediate symptom, the maintenance team will inherit a latent defect that may cause unexpected downtime months later.
One of the most important maintenance implications is that the fault may point to a missing or inadequate routine. For example, if a conveyor zone repeatedly accumulates debris because the cover plate does not seal correctly, then the maintenance schedule must include a check for that debris. In this way, the acceptance phase serves as the first opportunity to build the maintenance instructions from real evidence rather than from design assumptions.
Another implication is the need for spare parts planning. An investigation frequently reveals that a component has failed because it is operating outside its intended duty cycle. Replacing it with the same specification will only postpone the next failure. The maintenance plan should consider whether an alternative component, a different mounting position, or a change to the control logic would reduce the stress on the part.
Documentation is an ongoing maintenance implication. The findings from the fault investigation, including the evidence photographs, the root cause analysis, and the corrective action taken, should be filed in the equipment history record. This is not primarily a matter of compliance; it is a means of giving the next investigator a head start. Without a shared history, each fault investigation begins from zero.
Decision Boundaries: When to Stop and Escalate #
A disciplined fault investigation is one that knows when it has gone far enough and when it has crossed a boundary that requires a different level of authority. The investigator must be honest about the limits of their own knowledge and about the quality of the evidence in front of them.
The first boundary is drawn by lockout and safety rules. If the investigation reaches a point at which the cause of the fault is not yet clear but the next step would require momentary energisation of the system, that step must be governed entirely by the site permit-to-work policy and the OEM procedures for controlled testing. No investigator should make a unilateral decision to remove a safety guard or to bypass an interlock, even briefly. Such actions are never justified by a desire to speed up the investigation.
The second boundary is drawn by the limits of diagnostic equipment. If the fault requires a measurement technique that is not available on site, such as high-frequency signal analysis or thermal imaging, the investigator should pause and request specialist support rather than guess. A careful guess is still a guess, and in a safety-related circuit, a guess can be dangerous.
The third boundary is drawn by the difference between a root cause and a contributing cause. If the investigation identifies a root cause, such as a failed bearing, but also discovers that the bearing was loaded incorrectly because of a design flaw in the mounting bracket, then the investigator has actually found two issues. The design flaw is beyond the scope of a simple repair and must be escalated to the OEM or the design department. Stopping the investigation at the bearing replacement would leave the design flaw in place.
Escalation should also occur when the same fault repeats after multiple interventions. This is a clear sign that the investigation model is wrong. Repeating the same repair is not a further test of that model; it is a waste of time and a hazard. The investigator should escalate to a more senior engineer or to a team with access to deeper system knowledge.
Finally, the investigation must be brought to a defined close. The operating team must be informed that the system is ready for a controlled restart, or the system must be left in a documented non-operational state if it is not ready. Leaving a system in a half-repaired state without notifying the shift leader is a serious error in commissioning practice. The handover to operations must include an explicit statement of what was found, what was corrected, what was not, and what trials remain necessary.
Key Takeaways #
- Safe fault investigation is a controlled discipline that prioritises evidence preservation and boundary setting over the speed of recovery, with site lockout procedures and OEM documentation always taking precedence.
- A single visible fault in a warehouse system is often the symptom of an interaction between mechanical, electrical, and software subsystems, so the investigation must extend beyond the component that is in view.
- The evidence record, including photographs, verbatim fault messages, time stamps, and witness statements, must be collected before any adjustment is made.
- Intermittent faults and multiple simultaneous faults should be approached by searching for the shared condition or shared resource rather than by treating each symptom as an independent problem.
- Common interpretation errors include blaming the most recent change, assuming a component has only been in its current state, and drawing conclusions from identical components that are not identically stressed.
- Findings from the acceptance phase must feed directly into maintenance schedules and spare parts strategy, so that latent defects are not passed into routine operations.
- Escalation is the correct action when the investigation exceeds available diagnostic tools, when the root cause exposes a design issue, or when the same fault repeats after multiple repairs.
- The close of an investigation must produce a clear handover statement that informs operations what may be restarted, what remains pending, and what trials are still required.