Key takeaways
- Low volume makes percentage productivity claims misleading on small floors.
- Pass or fail on task completion, exception rate, and supervisor minutes per shift.
- Labor availability and backup coverage matter more than headcount saved on day one.
- Quality checks beat vanity metrics when a missed aisle is visible to customers.
- Expansion criteria should name the next zone, not a vague fleet goal.
Why do small facilities need different pilot success criteria?
A robot pilot at a compact grocery, clinic, or branch warehouse rarely produces the dramatic percentage gains you see in a million square foot distribution center. One autonomous scrubber on ten aisles can look modest on a spreadsheet even when it reliably covers the only route that matters after close.
Small sites still win when the robot removes a task nobody could staff, keeps quality steady on thin crews, and proves supervision load stays manageable. Success criteria should reflect that reality instead of copying enterprise key performance indicators built for high volume.
The right framework asks whether the unit completes defined tasks on schedule, whether exceptions stay rare, and whether the building can absorb the robot without reshaping the whole operation on day three.
What should task reliability measure during the pilot?
Task reliability starts with a written task list, not a vague promise of automation. Name the zone, the window, and the completion standard. An overnight scrub pass might require full aisle coverage with zero uncleared stops. A delivery loop might require tray arrival within a set minute band on three consecutive nights.
Track completed tasks versus planned tasks, not square feet alone. Record every stop with a reason code: obstacle, battery, human hold, software fault. A small site can tolerate fewer total hours if nearly every planned task finishes without rescue.
According to the U.S. Bureau of Labor Statistics Job Openings and Labor Turnover Survey, transportation, warehousing, and utilities posted a 4.0 percent average monthly total separations rate in 2025. Small logistics bays feel that churn in lost know-how, which makes repeatable robot tasks more valuable than a one-week productivity spike.
- Planned tasks completed without manual takeover
- Documented stop reasons per shift
- Charge cycles that fit the published schedule
- Recovery time after a fault below an agreed threshold

How do labor availability and supervision factor in?
Small facilities often run with a supervisor who is also the closer, the opener, or the maintenance contact. Pilot success includes supervisor minutes per shift spent on the robot: start checks, exception clears, and end-of-run sign-off.
If the robot adds more daily supervision than it removes walking time, the pilot fails even when the machine runs. Measure who must be on call if the unit stops at 2 a.m. and whether backup coverage exists under your service plan.
Labor availability also means honesty about unfilled work. A scrubber that backstops an overnight pass nobody would hire for is a win. A delivery loop that still needs a runner beside it is a scope problem, not a robot failure.
Why does quality consistency beat big percentage claims?

Customers and auditors notice missed spots on a small floor faster than on a cavernous warehouse. Quality criteria should include visual standards your team already uses, plus any check you trust such as wet floor appearance, squeegee lines, or tray delivery accuracy.
In the Phoenix Mesa Chandler metro, BLS May 2024 data lists building and grounds cleaning and maintenance workers at a mean of about $19.00 per hour. That wage band is not the whole story, but it shows why consistent overnight coverage can matter as much as raw speed when hiring stays tight.
Score the pilot on consecutive nights that meet the quality bar, not on a single hero run. Two weeks of boring consistency beats one perfect demo followed by drift.
What space and schedule constraints should gate the pilot?
Small buildings hide friction in narrow aisles, single docks, shared corridors, and one elevator bank. Success criteria should list physical constraints up front: minimum aisle width, ramp slope, door hold times, and hours when humans share the path.
Schedule constraints matter equally. A clinic that cannot close a wing, or a retail box that reopens early, needs explicit quiet hours for mapping and scrubbing. Pass the pilot only if the robot keeps those windows without constant reprogramming.
If the unit needs daily map edits to avoid people, capture that as a supervision cost in the same scorecard.
When has the pilot earned expansion?
Expansion should mean the next defined zone or the next task type, not a vague fleet purchase. Name the adjacent aisle, the second delivery loop, or the back room that opens only after the first route stays stable for thirty consecutive planned runs.
Finance will ask what changes on the monthly invoice. Tie expansion to measurable workload: another night pass, another patrol loop, or another tote lane that currently burns hourly labor.
Service Robot Co. often structures robot rental monthly or month to month lease paths so a small site can add capacity after the scorecard passes without repeating a full capital debate. One partner still owns select, finance, deploy, and service with a nationwide network of regional service engineers.
How should you document the scorecard for stakeholders?

Keep one page visible to operations and finance. List the task, the window, the reliability target, the quality check, supervision minutes, and the expansion trigger. Update it weekly during the pilot, not from memory at the end.
Photos of the floor or simple checklists beat narrative summaries. Exceptions belong in a log with dates so you can see whether a fault repeats in the same aisle or at the same door.
A passed pilot ends with a signed go or no-go on the next zone. A failed pilot ends with a clear reason: scope too big, space too tight, supervision too heavy, or service response too slow. All three outcomes are useful if they are honest.



