Key takeaways
- Write acceptance criteria before the robot ships, not after the first bad night.
- Benchmark on your aisles, doors, Wi-Fi, and peak traffic, not a demo lane.
- Score autonomy, productivity, safety, recovery, consumables, and operator touches in one rubric.
- Run enough cycles to see battery, comms, and repeat-grasp variation surface.
- A paid pilot should convert only when SAT-style evidence is already on file.
What does benchmarking a robot before a paid pilot actually mean?
Benchmarking is not a polished demo on a taped floor. It is a short, structured proof on your real building, with your loads, your obstacles, and the people who will share the space. You agree up front what good looks like, how many runs count, and what happens when the robot stalls, loses signal, or needs a human reset.
A paid pilot should extend work you already measured, not discover gaps you never wrote down. Procurement teams have used factory and site acceptance language for decades in process automation. IEC 62381 describes formal factory acceptance, site acceptance, and site integration testing with written plans, punch lists, and documented results. Service robots deserve the same discipline even when the purchase is a monthly rental instead of a capital project.
If you cannot explain pass/fail in one page, you are not ready to commit cash beyond a refundable assessment week.
Why do paid pilots fail without a written acceptance rubric?
Most disappointments trace back to vague promises. Words like reliable navigation or safe around staff mean nothing until you define the test lane, the success rate, and the allowed number of human interventions per shift.
Vendor labs hide the variables that bite on day three at your site: glossy patches, floor joints, dock plates, pallet overhang, reflective glass, and radio dead zones behind steel racking. CEN and ISO guidance on service robot performance stresses that test methods must match the application subcategory, and that performance criteria are not interchangeable with safety sign-off. Benchmarking separates what the machine can do in your building from what your safety review must still approve.
A rubric protects both sides. The buyer gets evidence tied to contract language. The integrator gets a fair shot to tune maps, speeds, and recovery rules before you scale fleet count.
Which six measures belong in every pre-pilot scorecard?
Treat the scorecard like a balanced report card, not a single shiny metric. Autonomy covers completion without unplanned help. Productivity covers area per hour, cycle time, or trips per shift against your baseline labor plan. Safety covers stops, separation, and documented responses when someone crosses the path.
Recovery covers what happens after a lost localization, a blocked aisle, or a low battery event. Consumables covers water, detergent, pads, bags, or charging duty cycles you will actually buy. Operator effort counts escorts, manual resets, and minutes spent babysitting per night.
Weight the categories for your job. A scrubber pilot might weight coverage and dry time highest. A delivery loop might weight door timing and elevator handoffs. An AMR tote move might weight dock congestion and WMS handshake errors.
- Autonomy: percent of planned missions finished without unplanned intervention.
- Productivity: throughput versus your current manual method on the same window.
- Safety: compliant stops and documented reactions to people and forklifts.
- Recovery: time to safe state and steps to resume after injected faults.
- Consumables: fill intervals, pad wear, and battery margin on longest route.
- Operator effort: counted touches per shift with names attached to each task.
How should you design routes on real floors and real loads?
Start from the worst honest shift, not the quiet Tuesday you wish you had. Map the full loop including turns you hate, pinch points, and the doorway everyone props open. Use the same pallets, totes, carts, or spill types you see in production, including beat-up versions with shifted loads.
Mark observation zones where a supervisor will stand without helping unless safety rules require it. Photograph floor defects you already know about so later disputes reference the same cracks and seams.
Run mixed traffic if forklifts, guests, or students share the lane during part of the day. Note whether the robot slows, reroutes, or requests help, and whether that behavior is acceptable for go live.

How many cycles are enough to trust the numbers?

One perfect lap proves marketing, not operations. Plan enough repetitions to heat batteries, stress Wi-Fi handoffs, and surface intermittent perception misses. Many integrators scope fifty consecutive task attempts as a starting point for success rate when the task is discrete and repeatable, then extend hours for cleaning coverage or patrol duration.
Inject controlled faults: block the path once, dim the lights if night work matters, and simulate a charger offline. Record whether the unit enters a safe state, pages a human, or recovers alone within your time limit.
Document environmental bounds. If results hold only between empty and half full occupancy, say so. If rain tracked in from a lobby mat changes stopping distance, capture that too.
What safety evidence should you collect before expanding scope?
Benchmarking is not a substitute for your risk assessment, but it should feed it. Note separation behavior when people approach from blind corners. For collaborative workcells, separation monitoring tests belong in the safety file even when this article focuses on mobile service robots.
Log every estop, bumper event, and software pause with timestamp and cause code. Compare against your permit to operate rules for fire lanes, elevator use, and restricted zones.
Keep video aligned to route IDs so insurers and corporate safety can replay incidents without guessing which building wing you meant.

How do consumables and maintenance show up during a short benchmark?
A scrubber that drains tanks twice per aisle may still fail your labor model even if navigation is flawless. Measure refill time, dump time, and pad changes across the full route, not just the easy half.
Battery plots matter for patrol and delivery loops. If the unit enters derate late in the shift, your night coverage math is wrong.
Ask what remote triage can fix versus what still needs a technician on site. That split drives whether your pilot length matches real maintenance calendars.
Who signs acceptance, and what triggers a paid pilot?
Name signatories before day one: facilities, safety, IT for wireless, and the manager who owns labor hours. Each signs the category they control.
Convert to a paid pilot only when mandatory thresholds pass on your site data. Optional stretch goals can wait. If autonomy passes but productivity misses, negotiate scope reduction instead of paying for fleet scale you cannot use.
Attach raw logs, map revisions, and exception codes to the signoff PDF so renewals do not restart from zero.
Where does Service Robot Co. fit a benchmark week?
We arrive vendor neutral with a draft rubric tailored to your building type, then tune it with your safety and operations leads before the first run. Routes are mapped on your floors, finance paths stay open for rental or purchase, and our nationwide network of regional service engineers backs remote triage plus on site dispatch when a benchmark exposes a hardware limit.
A free site assessment can narrow robot type and charger placement before you spend on a multi week pilot. If the benchmark fails, you still keep the documentation for the next vendor conversation instead of a vague memory of a trade show demo.



