Technician Diagnostic Checklists: Data Signals and Condition Monitoring #
In automated warehouse systems, the boundary between a mechanical fault, an electrical fault, and a control logic fault is rarely clear at first observation. A stopped conveyor, a missed barcode read, or an unseated carton on a sorter is the visible outcome of a chain of interactions. A technician diagnostic checklist is a structured method for capturing both transient data signals and persistent condition evidence before the machine is reset and the evidence is lost. This article explains how to design such checklists for use by maintenance technicians, controls teams, and warehouse operators, with emphasis on gathering data, interpreting what the data actually means, and using documented evidence to reduce repeat faults.
The Role of Diagnostic Checklists in Warehouse Automation #
Checklists are not substitutes for experience, nor are they rigid scripts that ignore live conditions. They are a way of standardising the first minutes after a fault is reported. In a typical automated warehouse, the technician arrives after the HMI has shown an alarm, the operator has already tried to reset the line once, and the fault may or may not reproduce. Without a checklist, the natural tendency is to focus on the most recent alarm code and assume it identifies the root cause. That assumption is often wrong.
A diagnostic checklist achieves three things. First, it forces a consistent observation sequence: what changed, when it changed, where the change occurred in the physical system, and what data was present at the moment of the event. Second, it separates live signal data from condition monitoring evidence. A sensor output is a point-in-time electrical state; a bearing temperature trend is a condition that develops over hours. Third, it creates a record that can be compared with previous faults. Over time, that record reveals patterns that single events never show.
The checklist should be considered part of the machine’s maintenance documentation, not a standalone form. It must align with site procedures, lockout requirements, OEM documentation, and competent engineering judgment. No checklist overrides the mandatory rules for safe isolation of energy sources.
Data Signals as First-Response Evidence #
Data signals are the earliest and most objective evidence available after a fault. In modern controllers, the PLC or programmable automation controller logs inputs, outputs, and internal variables at the time of an alarm. These logs are the starting point for diagnosis. However, they are often treated as conclusive when they should be treated as directional.
A signal such as “conveyor 12 overload” tells you that the motor protection device tripped. It does not tell you whether the root cause was a jammed carton, a seized bearing, a brake that did not release, a misadjusted sensor causing repeated start-stop pulses, or a soft starter that degraded over time. Each possible root cause has a different remedy, and each has a different prevention strategy. The signal narrowed the problem area; the checklist broadens the evidence base.
Signal Categories and Their Meaning #
For the purpose of diagnostic checklists, data signals can be grouped into six categories. Each category should be captured in a consistent format on the checklist.
- Discrete inputs: Photoelectric sensors, proximity switches, limit switches, and mechanical interlocks. These indicate presence, position, or state. A discrete input that is stuck high or low is a signal failure, not necessarily a mechanical failure.
- Discrete outputs: Commands to contactors, relays, solenoid valves, and indicator lights. An output that is commanded on but produces no downstream effect points to wiring, interposing relay, or actuator failure.
- Analogue signals: Encoder counts, position feedback, load cell readings, pressure transmitter outputs, and analogue temperature inputs. These provide magnitude and direction, and trend data is often more valuable than instantaneous values.
- Drive and motor data: Current draw, DC bus voltage, motor frequency, torque reference, and fault codes from variable frequency drives (VFDs) and servo drives. These are monitored by the drive itself but should be cross-checked against mechanical behaviour.
- Network and communication signals: I/O adapter status, switch port statistics, packet loss counters, and diagnostic messages from AS-i, Profinet, EtherNet/IP, or similar plant networks. Intermittent faults are frequently caused by network interference, grounding problems, or damaged cables.
- System-level status: Mode changed, reset commands, operator acknowledge events, and sequence interlocks. These are often the only way to know what the operator did before the fault was reported.
Each category needs a different collection method. Discrete and analogue signals are usually captured by the controller at scan rates. Drive data is captured by the drive’s internal log. Network signals are captured by switches or diagnostic tools. System-level status is captured by the HMI or SCADA historian. A good checklist indicates which source to query first for each fault class.
Condition Monitoring Parameters for Mechanical Systems #
Condition monitoring differs from data signals in an important way. Data signals respond to events that already exist; condition monitoring captures the slow drift that precedes failure. In warehouse automation, the most common monitored condition is mechanical health of rotating elements: conveyor rollers, gearboxes, motors, sorter cross-belts, palletiser clamps, and AS/RS mast mechanisms.
Condition monitoring evidence does not always require an installed sensor. A technician’s own senses are valid condition monitoring instruments, provided the observations are disciplined. The checklist should include spaces for measured values such as:
- Surface temperature of motor housings, gearbox casings, and bearing caps, recorded using an infrared thermometer or contact probe.
- Vibration severity readings if a handheld vibration meter is available, or a descriptive note on abnormal noise level and character.
- Audible changes: ticking, grinding, whining, or repetitive thumping that changes with line speed.
- Visual evidence: leaked oil, dust accumulation, scoring on shafts, fretting corrosion around bolts, and pitting on sprockets or belts.
- Speed and position consistency measured through encoder counts during repeated cycles.
The value of this evidence depends on trend comparison, not single readings. A gearbox running at 50 °C in summer and 55 °C in winter may be normal; the same gearbox running at 70 °C one month after the previous inspection is a warning. The checklist should therefore include a field for previous recorded values, even if the technician has to request them from the CMMS (computerised maintenance management system) or look them up on the lubrication schedule.
Building the Technician Diagnostic Checklist #
A checklist for warehouse automation diagnostics should be organised around the sequence of the technician’s work, not around the system hierarchy. The design intent is that a technician can follow it in the field without needing to read the full electrical schematic or the OEM manual first. Three phases cover the majority of diagnostic activity.
Pre-Trip: Before Running the System #
This phase begins immediately after the fault is reported, before any energy is applied or removed. The technician records the operating state at the time of failure, reviews the HMI alarm history, and identifies the exact machine zone. The pre-trip phase is critical because the machine is still in a safe, stopped state. Evidence that is easy to collect now may be impossible to collect after a reset.
Actions in the pre-trip phase include checking for blocked sensors, pinched cables, loose fasteners, and pooling oil around the stopped equipment. The logic here is simple: if a photoelectric sensor is blocked by a torn piece of stretch film, resetting the line will reproduce the fault immediately. The technician should also check the orientation of the product at the point of the fault. A tilted carton indicates a handling problem upstream, not a sensor problem at the point of detection.
Running-State Observations #
In this phase, the system is run under controlled conditions, usually in manual or jog mode and within the limits allowed by site procedures. The technician observes whether the fault reproduces, notes the speed at which it occurs, and watches the interaction between moving product and sensing zones. The running-state phase is where data signals and condition evidence are linked.
For example, if a conveyor section trips its overload relay only when a specific pallet type passes through a curve, the technician should note the shape of the pallet, the gap between the pallet and the conveyor side rails, and the speed at the time of the trip. These observations give context to the PLC alarm. Without them, the overload trip will be attributed to a random jam, and the underlying cause—in this case a bent pallet rubbing against the rail—will remain in the system for weeks.
Running-state observations should have a time limit. If the technician cannot reproduce the fault within a reasonable number of cycles, the diagnostic activity should move to a planned inspection and not continue indefinitely in production time. The checklist should include a decision point: “Fault reproduced: yes/no. If no, proceed to intermittent fault protocol.”
Post-Trip and Intermittent Fault Observations #
After the fault has been reproduced once, the technician should step back and observe again, this time with a focus on what was happening on the line immediately before the trip. A common source of repeat faults is the interaction of two components that are individually healthy but collectively unstable. For instance, an encoder signal may be noisy due to a damaged cable every time a certain sorter tray passes a fixed point. The sorter operates normally at low speed but loses its position reference at full speed. The PLC sees a position error and stops the sorter just after the noisy section.
Intermittent faults require a different evidence collection strategy. The checklist should prompt the technician to capture the following before the next occurrence, rather than waiting for the alarm history to grow:
- Exact time and duration of each alarm event.
- Which I/O module, network node, and terminal number were active in the scan before the fault.
- Whether the fault follows the operator shift, the product type, the ambient temperature, or the line speed.
- Whether the fault first occurred after recent mechanical work, lubrication, or wiring changes.
A Practical Diagnostic Table for Signal-Condition Matching #
The table below shows common symptom-signal combinations, the condition monitoring evidence to collect, and the first-line checks that link the two. It is not a complete fault matrix; it is a template the site maintenance team should adapt to its equipment.
| Symptom | Typical Data Signal | Condition Evidence to Collect | Likely Interactions | First-Line Checks |
|---|---|---|---|---|
| Conveyor zone stalls under load | Motor overload trip; drive fault code | Motor housing temperature; belt or chain tension; roller resistance when spun by hand | Mechanical overload reflected in electrical current; weak drive braking mixes; product jam upstream | Check for jammed product in the zone; verify the motor cooling fan is clear; compare actual current against drive nameplate rating |
| AS/RS crane stops at the same rack position | Position error; encoder count mismatch | Rail contamination; mast deflection at that height; loose encoder coupling | Encoder slipping at a specific height due to rail wear; thermal expansion affecting sensor mounting | Slow-move the crane to the position and inspect the rail joint; verify encoder coupling grub screws are tight |
| Sorter misses a barcode read intermittently | No-read alarm; sensor output low pulse width | Reader illumination condition; lens contamination; label quality on the carton | Product flapping on conveyor; ambient light interference; degraded label print contrast | Confirm the sensor trigger is stable; clean the lens; test with a known-good label |
| Cross-belt sorter repeatedly rejects one tray type | Drive overcurrent for one tray; communication timeout on that node | Vibration level of the tray; belt tracking under that tray; connector seating on the moving deck | Failing bearing on one tray increases current; intermittent connector contact appears as a network timeout | Manually push the tray section through a full rotation and listen for noise; reseat all connectors on that tray |
| Palletiser clamp loses position | Analogue position feedback out of tolerance | Hydraulic or pneumatic pressure at the actuator; clamp pad wear; sensor mounting bracket looseness | Sensor feedback decays due to a failing power supply; actuator internal leakage changes speed and stroke | Compare the displayed position against a known physical reference; check supply pressure; recalibrate the sensor if authorised |
Common Interpretation Errors #
Several recurring errors undermine the effectiveness of signal-based diagnosis. The most common is treating the alarm as the root cause. An HMI alarm is a designed response to a condition, not the reason the condition exists. If the PLC is programmed to issue an “air pressure low” alarm when the pressure switch changes state, the alarm is a description of the switch position. The root cause may be a faulty pressure switch, a failed compressor, a leaking pipe, or a blocked filter. Every one of these produces the same alarm.
A second error is ignoring the sequence of alarms. Alarm logs are often read as independent events, but a lead alarm may be irrelevant. A technician who looks only at the last alarm will miss the fact that a supplier conveyor was already stopping and restarting six times a minute due to a sensor delay, and that the final overload trip was simply the consequence of repeated inrush currents.
A third error is over-reliance on the most recent reset. In automated warehouses, operators often reset a fault several times before calling a technician. Each reset may clear the data log of the exact values that were present at the first occurrence. The checklist should instruct technicians to verify whether the fault is fresh or has been reset before their arrival, and to record the number of resets. This number is valuable condition evidence: it suggests that the system is responding to an unstable condition, not a one-off event.
A fourth error is confusing signal quality with signal level. A digital signal that is present but noisy—for example a 24 V DC signal dropping to 19 V under load—may still cause a PLC input to flicker. The PLC sees a change of state that never actually occurred at the sensor. This issue is invisible on the HMI and visible only on a handheld meter or an oscilloscope during a controlled test. Condition monitoring of supply voltages at the sensor head, and checks of connector backshells and grounding continuity, are the relevant remedies.
Failure Coding and Repeat-Fault Reduction #
The purpose of collecting diagnostic evidence is not only to fix the current fault but to prevent the next one. This requires failure coding that is granular enough to distinguish between failure modes. A code such as “sensor replacement” is not informative if it is applied to every case where a photoelectric sensor was involved in a fault. The sensor may have failed due to contamination, due to a damaged cable, due to vibration loosening its mounting bracket, or due to a damaged reflector. Each of these causes has a different corrective action and a different prevention strategy.
In practice, the technician should enter at least three pieces of information into the maintenance record after the fault is resolved:
- The observed symptom, written in the language of the operator, for example “Cartons jam at the merge point between line 2 and line 3.”
- The physical cause confirmed by inspection or testing, for example “Freely rotating side guide moved 4 mm due to a loose M8 bolt.”
- The prevention action taken, for example “Applied thread-locking compound to the bolt and added the side guide alignment check to the monthly checklist.”
This simple classification makes repeat-fault reduction possible. The CMMS can then report whether the same physical cause is recurring, regardless of which alarm code was triggered. Without this granularity, two identical faults might be logged under different alarm codes simply because the HMI message differed, and the pattern would never be visible.
Spares Strategy and Evidence-Driven Replacement #
Diagnostic evidence also guides the spares strategy. When a technician has recorded the physical cause of a confirmed failure—not just the alarm—the organisation can determine which spare parts are consumed repeatedly, which are replaced prematurely, and which are never used. Premature replacement occurs when a part is swapped on the assumption that it is faulty, but the real cause is a wiring fault or a mechanical misalignment. The removed part is then tested as “good” and may be discarded or returned to stock, introducing uncertainty into the spare parts inventory.
The checklist should therefore include a verification step after replacement: before closing the work order, the technician should confirm that the newly installed part operates at the expected signal level and that the original condition has been resolved. This step is often skipped in the pressure of production demand. However, its absence is a leading cause of recurring faults and of spare parts being blamed unfairly for repeat failures.
Evidence-based replacement also applies to condition monitoring outcomes. If a bearing temperature has been trending upward over several months and the checklist shows a steady increase in vibration level, replacing the bearing during a planned maintenance window is justified. Replacing the same bearing because the sorter stopped once is not necessarily justified, unless the condition evidence supports a relationship between the bearing state and the stop event.
Decision Boundaries and Escalation #
Technician diagnostic checklists must define clear decision boundaries, otherwise they become an excuse for indefinite troubleshooting in production time. A reasonable boundary is: if the fault is reproduced and the physical cause is identified to the level of a specific component or adjustment, the technician can proceed with the corrective action. If the fault cannot be reproduced after a specified number of cycles, or if the data signals point to a conflict between multiple controllers, the diagnostic activity should be escalated to the controls engineering or OEM support level.
Escalation is not a failure of the diagnostic process. It is the correct outcome when the checklist has exhausted the evidence that can logically be extracted without deeper system knowledge. The escalated report should include the checklist record itself—what was observed, what was measured, what was replaced, and what changed after each action. That record is far more valuable to the controls team than a verbal summary of “we tried a new sensor and it still faults.”
The decision boundary also applies to safety-related faults. If a signal indicates a safety device was activated, the technician must not bypass or override it. The only valid response is to follow site procedures for inspecting and resetting the safety circuit, and to refer to the OEM documentation for the system-specific behaviour. The checklist can note that a safety device activation has occurred, but the remedy is governed by site rules, not by the diagnostic checklist.
Finally, each decision boundary must respect the original equipment manufacturer’s design intent. Any modification to the machine, the control logic, or the wiring that goes beyond the OEM’s stated adjustment range requires authorisation from an engineer with competence for the system. The checklist is a tool for diagnosis and record-keeping, not an authorisation for engineering change.
Key Takeaways #
- Diagnostic checklists preserve early evidence—data signals and condition monitoring—before fault reset routines erase it.
- Treat HMI alarms as directional clues, not root-cause statements; the same alarm can follow from mechanical, electrical, network, or operator-related causes.
- Separate live point-in-time data signals from trend-based condition evidence; both are needed for a reliable diagnosis.
- Capture specific physical evidence in the pre-trip phase, when the machine is stopped and safe to inspect.
- Use an adaptive diagnostic structure: stop the fault, observe the running state under controlled conditions, and follow a separate protocol for intermittent faults.
- Record failures at the level of physical cause and prevention action
Related Pearl Gateway Guides #