Site Acceptance Testing (SAT) in warehouse automation is a formal, time-boxed evaluation of an integrated material handling system at its installed location, using the actual building, electrical supply, control network, software configuration, and, ideally, the operators who will run the system daily. It is a distinct stage between Factory Acceptance Testing (FAT), which proves individual machines under controlled factory conditions, and steady-state production, which proves commercial output over weeks. SAT asks a narrow but critical question: does the installed system, acting as a whole, meet the agreed operational requirements under agreed boundary conditions? Because the answer shapes contractual handover, warranty periods, spare parts holding, and staffing plans, the selection of what to test, how to test it, and what to do with the results deserves deliberate engineering attention. This article outlines the criteria for choosing a SAT scope, the boundaries of what a SAT can validly prove, and the interpretation rules that keep SAT evidence useful for the entire life of the installation.
Before any SAT activity, site procedures, lockout requirements, OEM documentation and competent engineering judgment take priority over any generic guidance, including this article. SAT is not an occasion for improvisation; it is a disciplined evidence-gathering exercise that must respect the physical and administrative safety controls of the warehouse.
The Role of SAT in an Automation Lifecycle #
A warehouse automation project does not move from design to operation in one step. It passes through staged validation, and each stage has a different purpose. FAT confirms that a machine or skid is built correctly. Installation and commissioning confirm that the machine is connected, powered, and safe. SAT confirms that the machine plus its neighbours plus the software plus the network plus the operators form a coherent system. The distinction matters because most performance failures in automated warehouses are emergent: no single component is defective, yet the system behaves poorly because of an interaction between conveyor speed, order release logic, and upstream scanner performance.
SAT is best treated as a baseline event, not a final exam. It produces a calibrated snapshot of system behaviour at a known moment in time, under known conditions, with a defined set of inputs. That snapshot becomes the reference point for later change control. If a warehouse operator can look back at the SAT data after two years of operation and compare current throughput, error rates and recovery patterns against that baseline, the SAT has delivered its real value. If the SAT data exists only as a signed-off certificate in a project folder, the testing exercise was administratively complete but technically wasted.
Selection Criteria: When a Full SAT Is Justified #
Not every system modification requires a full site acceptance test. Running a full SAT is expensive in time, order queue capacity, and staffing. The selection of a SAT scope should balance the cost of testing against the risk of an unforeseen interaction reaching production. Three criteria dominate the decision.
Criticality and Novelty #
A full SAT is justified when the installation includes new technology, a new control architecture, or a product handling principle that the site has not used before. Novelty reduces the validity of past experience. If the warehouse is commissioning a new type of sortation induction or a new robotic palletising cell, the interaction behaviour cannot be inferred from an older system. Similarly, higher criticality justifies a broader SAT. If the system under test is the sole outbound path for a large-format warehouse, the cost of a week of testing is small compared with the cost of a failed ramp-up. If the system is a low-throughput returns line with ample manual backup, a lighter verification may be reasonable.
Integration Complexity #
Integration complexity is not measured by the number of machines but by the number of asynchronous control domains that must exchange state. A conveyor that moves cartons from a single infeed to a single outfeed may have only one PLC and one high-level order release; its integration risk is low. A system with several induction points, a crossbelt sorter, a robotic depalletiser, a stretch wrapper, and a warehouse control system that prioritises orders dynamically has high integration complexity. Each interchange point is a potential source of lost time. For such systems, the SAT must cover the interchanges, not just the endpoints.
Contractual and Operational Timing #
The timing of the SAT relative to business seasonality is a legitimate selection criterion. A SAT conducted during a peak shipping window will almost certainly be polluted by expedited orders, an abnormal SKU mix, and staffing levels that do not match the steady-state plan. Conversely, a SAT conducted too early, before operators have been trained or before critical spares are on site, delivers results that are not representative of sustainable operation. The decision to run the SAT should therefore include a readiness review: is the skill level of the operating team, the stock profile of the test orders, and the availability of support engineers aligned with the conditions the SAT is supposed to certify?
Application Boundaries: What SAT Cannot Reasonably Prove #
An SAT is bounded in time, in input variety, and in system configuration. Recognising these boundaries prevents the evidence from being over-interpreted. The most common boundary errors relate to the following limitations.
First, SAT cannot prove long-term reliability. A two-day test will not reveal abrasive wear on a belt, gradual misalignment of a photocell, or the slow deterioration of a battery pack under repeated partial charging. SAT evidence can only establish that the system starts and operates with acceptable performance at the beginning of its life, not that it will do so in six months. Second, SAT cannot prove performance under an unanticipated SKU mix. Any fixed test plan defines a finite set of carton sizes, weights, and orientations. A warehouse that later introduces an irregularly shaped product, a more fragile package, or a heavier box is operating outside the validated envelope and should expect behaviour that the SAT did not observe.
Third, SAT cannot certify software under all future revisions. The SAT applies to a specific software version, a specific parameter set, and a specific database state. If the controls team subsequently changes the order release algorithm, adjusts a timer, or adds a new scanner diagnostic, the SAT evidence becomes stale for the modified scope. Fourth, SAT does not validate operator ergonomics, morale, or safe practice under sustained pressure. The operators who take part in a carefully planned SAT are not experiencing the same cognitive load as a shift that has run at capacity for five consecutive days.
Component Interactions That Shape SAT Design #
A useful SAT design starts by listing the component interactions that are most likely to cause emergent failure. These interactions are often invisible in individual machine documentation because each vendor describes its own equipment, not the handshakes between zones.
One high-value interaction is the relationship between conveyor speed, accumulation capacity, and sortation induction timing. If the infeed conveyor runs faster than the sorter can accept, jams at the induction point are inevitable, even though both machines meet their individual specifications. The SAT must measure gap behaviour at the induction point under a realistic mix of carton lengths. A second interaction is the control loop between the warehouse control system (WCS) and the mechanical system. If the WCS releases orders in bursts that overwhelm the buffer, the system will experience alternating starvation and oversaturation. The SAT should record release timestamps alongside zone occupancy to see whether the control algorithm produces a smooth flow.
Battery charging strategies create a third interaction class. Automated mobile robots, electric monorail conveyors, and automated guided vehicles share a constrained charging infrastructure. The interaction between scheduling software, charge thresholds, and actual machine power consumption determines whether robots reach their next mission or interrupt traffic. The SAT should run at least one full charge-discharge cycle per vehicle type, under a workload that matches the expected shift profile. Finally, the interaction between the warehouse management system (WMS) and the WCS deserves explicit testing. Wrong location assignments, inconsistent barcode data, and duplicate orders propagate from the higher-level system into the physical flow. The SAT should deliberately include a small percentage of intentionally imperfect data to observe recovery behaviour.
A Diagnostic Table for SAT Observation #
The table below lists observable symptoms during a SAT, the interaction that often lies behind the symptom, the evidence that should be recorded, and the boundary that should stop the tester from drawing an overly broad conclusion.
| Observed Condition | Component Interaction to Inspect | Evidence to Record | Interpretation Boundary |
|---|---|---|---|
| Throughput drops after the first 45 minutes | Induction backpressure, conveyor speed mismatch, scanner re-read loops, WCS order release throttling | Per-zone active and idle timestamps, order release log, scanner re-read count by hour | Does not prove long-term degradation; may indicate a buffer that is undersized for a specific SKU pattern |
| Intermittent jams at a merge point | Package gap variance, divert timing, sensor calibration, speed of upstream and downstream zones | Jam count per hour, jam location by zone, photograph of each jam, measured gap before and after | Not a defect unless the same product profile produces a repeatable jam on multiple runs |
| Mobile robots pause at charging stations more often than expected | Battery state-of-charge threshold, charge schedule priority, mission release logic, vehicle assignments | Charge start and end times, miss rate for mission deadlines, state-of-charge histogram | May be temporary if the SAT order mix differs sharply from the planned production mix |
| Error code rate rises slowly through the day | PLC communication health, WES database connection, device timeout settings, network traffic | Error code frequency per device, timestamp of first occurrence, controller CPU load, network gap count | Could indicate a network or software issue rather than a mechanical fault; retest after parameter adjustment |
| Operator intervention is needed to clear misaligned cartons before induction | Pre-induction orientation, length/height scanner coverage, conveyor side-guide adjustment, carton twist | Intervention count, misalignment type, SKU of the affected carton, duration of each stop | Acceptable only if the intervention rate is below the level agreed in the acceptance criteria |
Evidence Collection: Moving Beyond Pass/Fail #
SAT evidence is only as good as its provenance. A single overall number, such as “the system achieved 1,200 cartons per hour,” hides too much. The evidence collection plan should capture time-stamped, per-zone data that allows a controls engineer to reconstruct the sequence of events after a failure. At a minimum, the plan should record order release times, arrival times at each major control point, outfeed or destination times, and every exception event such as jam, re-read, reject, or timeout.
The testing team should separate system performance from operator performance. If the SAT requires manual induction, the speed and consistency of the inducer directly affects throughput. That effect should be measured and reported separately. One approach is to run the SAT in two phases: first, a hands-off phase in which the system handles an automated, fixed sequence of orders, and second, a realistic phase with operators working at a planned pace. Comparing the two phases shows how much of the measured performance depends on human skill rather than machine logic.
Ramp-up and ramp-down should be recorded explicitly. Many systems perform well in steady state but poorly at the transition from idle to full flow, or during the transition from a heavy load to a light load. The SAT should include several start-stop cycles, because the ramp-up transient is often where WCS/PLC synchronisation errors appear. Evidence from these transients is more valuable for ramp-up planning than the steady-state average, because a production day is full of starts and stops.
Common Interpretation Errors #
The most damaging SAT interpretation error is comparing the measured result to the theoretical maximum speed listed on the equipment nameplate. Nameplate speeds are single-machine, no-queue, perfect-conditions values. A system that achieves 85 percent of the theoretical sum of its components in a realistic SAT may be performing very well. A system that achieves 98 percent of the theoretical sum may actually be under-stressed, with buffer zones starved and operators idle. The correct comparison is between the SAT result and the agreed operational requirement, not between the SAT result and the sales brochure.
A second error is excluding recovery periods from the calculation. Every automated system stops occasionally. The difference between a good system and a poor one is often the time it takes to recover, not the frequency of stoppages. If the SAT records a jam that takes twelve minutes to clear and the recovery process requires an operator to climb a ladder, open a guard, and physically remove a carton, that recovery time belongs in the performance calculation. Smoothing it out by restarting the timer after each stop conceals a major operating cost.
A third error is allowing undocumented manual intervention during the test. If a supervisor steps in to push a carton, adjust a sensor, or manually trigger a restart, the intervention must be logged. Otherwise, the SAT conclusion is based on a system that does not exist in normal operation. A fourth error is extrapolating from a narrow SKU set. A SAT that tests only uniform cartons cannot validate a system that will handle a high mix of parcel sizes. The acceptance criteria should require that the test order profile reflects the actual distribution of product sizes and weights, with representative proportions of edge cases.
Maintenance Implications and Change Control #
The SAT is not only a handover event; it is also the first complete reference point for the maintenance