Key takeaways
- Judge the fleet by accepted work, not distance traveled, hours powered on, or missions launched.
- Use six KPIs: mission success, interventions, utilization, productive coverage, queue time, and recovery duration.
- Freeze metric definitions and exclusions before comparing shifts, sites, robot types, or reporting periods.
- Treat the first 90 days as a controlled learning period with daily triage, weekly operational reviews, and monthly business reviews.
Which metrics show if a new fleet is working?
Track six measures during the first 90 days: mission success, human interventions, utilization, productive coverage, queue time, and recovery duration. Together, they show if robots finish useful work, how much human effort autonomy still consumes, and where the operation is losing time.
No single KPI can answer the question. High utilization can conceal repeated rescues. A strong mission success rate can coexist with poor coverage if the fleet receives too little work. Fast recovery can mask a queue that leaves robots idle before missions even begin.
NIST advises that robot performance measures should be driven by application requirements and interpreted in context. Its current measurement program examines six facets: performance, collaboration, agility, autonomy, safety, and ease of implementation. A first-90-day scorecard should follow the same logic by measuring the work system, not merely the machine.
What counts as a successful mission?
Mission success is the percentage of eligible missions completed within the agreed operating criteria. The numerator is accepted completions. The denominator is every mission that started and was expected to finish, with cancellations and test runs reported separately rather than quietly removed.
Completion alone is too permissive. A transport mission should arrive at the correct destination, with the correct load, inside its service window. A cleaning mission should finish its assigned area and meet the chosen quality check. A patrol or inspection mission should return usable records from every required checkpoint.
Publish first-pass success beside eventual success. A robot that completes after three staff assists has produced a different result from one that completes autonomously. First-pass mission success exposes that distinction before retries make the dashboard look healthier than the operation.
NIST analyzed 1,464 human-robot interaction papers published across seven years and found that extensive use of custom measures limited baseline comparison. The practical lesson is simple: write one fleet-wide definition of success and keep it stable for the entire review period.

How should interventions be counted?

Count every unplanned human action required to continue, correct, or complete a mission. Normalize the total as interventions per 100 eligible missions so a busy shift can be compared fairly with a quiet one.
Tag each intervention by cause, actor, severity, and labor time. Useful cause codes include blocked path, localization loss, door or elevator issue, load problem, depleted consumable, docking failure, unsafe site condition, software fault, and operator mistake. Keep planned cleaning, charging, replenishment, and inspection in a separate maintenance category.
The intervention rate measures autonomy burden. Intervention minutes measure the burden on people. Both matter. Ten quick acknowledgments and one 40-minute physical recovery produce the same event count, but they impose very different operating costs.
Trend remote recoveries separately from cases requiring someone to reach the robot. Remote triage may consume only a control-room task, while an on-site dispatch can interrupt a supervisor, technician, or frontline employee. That distinction also identifies which recoveries belong in training and which belong in robot maintenance service.
Why do utilization and productive coverage need separate lines?
Utilization is active mission time divided by scheduled service time. It reveals how much of the fleet's planned window is spent executing assigned work. Exclude planned maintenance and document any charging policy so the denominator stays consistent.
Productive coverage measures accepted output against the work that was genuinely available. Use accepted square feet for floor care, correctly delivered loads for material movement, completed service stops for delivery, or checkpoints returning usable data for patrol and inspection.
The distinction prevents empty activity from winning. A robot may drive for most of a shift because routes are inefficient, work is repeatedly retried, or assignments send it through congested areas. That raises utilization without increasing useful output.
The U.S. Department of Energy describes overall equipment effectiveness as availability multiplied by performance and quality. It notes that 100 percent would require uninterrupted operation at maximum speed while producing only good output. A fleet scorecard need not copy that manufacturing formula, but it should preserve its core discipline: time, pace, and acceptable work are different dimensions.
What do queue time and recovery duration reveal?
Queue time starts when a valid job becomes ready for automation and ends when a robot begins executing it. Record the median and the 90th percentile by shift, task, and site. Averages alone can hide a small number of severe delays that upset production or service windows.
Long queues do not automatically mean the fleet is too small. They can originate in poor job release rules, congested pickup points, unavailable elevators, charging conflicts, manual handoffs, or a task mix that assigns several urgent missions at once. Compare queue growth with utilization before adding capacity.
Recovery duration begins when a mission-affecting fault or blockage is detected and ends when the robot returns to eligible service. Split that interval into detection, acknowledgment, diagnosis, travel, repair, validation, and release whenever the data permits.
Report median recovery for the normal experience and the 90th percentile for the difficult tail. Then pair duration with reason codes. Faster acknowledgment will not cure recurring localization failures, and quicker repair will not fix a supervisor approval step that leaves recovered units waiting for release.

How do you establish a credible baseline?
Start with the process the fleet was hired to improve. Capture comparable manual output, demand, completion quality, queue delay, exceptions, and labor touch time before go-live. If a pre-launch baseline is unavailable, preserve the first stable operating period and label it honestly as an early fleet baseline.
Freeze a data dictionary before the first formal comparison. For every KPI, specify the event that starts the clock, the event that stops it, the denominator, exclusions, source system, owner, and treatment of missing records. User-canceled missions, training runs, emergency stops, and planned maintenance should never drift between categories.
Segment the baseline by site, shift, task, operating zone, and robot class. Pooling unlike work can create false improvement. A quiet overnight route should not erase delays from a live daytime route, and a short internal delivery should not be averaged with a multi-floor trip.
NIST distinguishes concrete measures from higher-level metrics and stresses that measures should support decisions. That means every scorecard line needs a response rule. If a movement would not change routing, training, site design, maintenance, staffing, or fleet configuration, question why it is being collected.
What should the 90-day review cadence look like?
During days 1 through 14, validate data capture and classify commissioning events. Do not treat mapping corrections, staff training, and acceptance testing as steady-state production. Keep them visible, but in their own phase.
From days 15 through 30, stabilize definitions and identify repeated intervention causes. During days 31 through 60, change one operating factor at a time where practical, such as a route, job-release rule, charging window, pickup layout, or escalation path. Preserve a change log so KPI movements have an explanation.
Use days 61 through 90 to test repeatability across normal demand, shifts, and operating conditions. Compare the latest period with the frozen baseline and show both counts and rates. A percentage based on twelve missions deserves less confidence than the same percentage based on several thousand.
Run daily exception triage for safety events, failed missions, and prolonged recoveries. Hold a weekly operating review for trends and corrective actions, then a monthly business review connecting accepted output to service levels, labor touches, and the original deployment case. NIST warns that target-driven measurement can distort outcomes, while badly timed measurement cycles can miss or misread important events.
A fleet program needs one accountable operating picture
Mixed fleets often scatter evidence across separate portals, service tickets, shift notes, and spreadsheets. Service Robot Co. gives U.S. businesses one lifecycle partner for robot deployment and integration, financing, training, fleet oversight, and service through a nationwide engineer network.
Its OEM-neutral approach also matters during measurement. The integrator can apply common business definitions across manufacturers while retaining the diagnostic detail needed for each unit. That supports multi-vendor robot fleet management without pretending unlike tasks should have identical targets.
Commercial robot rental, robot leasing for business, and monthly payment programs still need the same operating proof as a purchase. Maintenance included in a service robot rental does not remove recovery duration from the scorecard. It makes the ownership and escalation path clearer.
By day 90, management should be able to decide which tasks are ready to scale, which require process redesign, and which should stop. One partner and one number for remote triage, on-site dispatch, training, and service keeps that decision tied to the entire operating system instead of a hardware activity report.



