Why a 99 Percent Success Rate Still Fails Most Robot Deployments
Dexterity, not locomotion or model size, is physical AI's real bottleneck in 2026, because per-action success rates compound across a workflow; ANYbotics, Agility Robotics and Starship Technologies show how matching a robot's capability to the minimum a job requires, and building recovery into the system, is what actually clears that math.

A robot that nails a pick-and-place demo nine times in a row looks, to most people watching, like a solved problem. Run that same 99 percent per-step success rate across a 100-step warehouse shift instead of a nine-step stage demo, and the odds of finishing the whole shift without a single failure fall to roughly 37 percent. That gap between a flawless-looking demo and a workflow a buyer can actually deploy unattended is where the humanoid and industrial-robotics industry's real bottleneck sits in October 2026, and it is not the one most funding rounds are priced around.
The bottleneck is dexterity, defined not as finger count or actuator sophistication but as a closed-loop system of sensing, control, recovery and judgment that has to hold up across an entire task sequence rather than a single rehearsed motion. That framing comes into sharper focus against three companies now running dexterous or semi-dexterous systems in paying deployments rather than lab demos. One of them, ANYbotics, is still unfamiliar to most buyers outside industrial inspection: the Zurich, Switzerland company spun out of ETH Zurich's Robotic Systems Lab in 2016 and builds the ANYmal quadruped, a four-legged inspection robot that reads gauges, checks for leaks and flags anomalies on oil rigs, wind farms and steel plants without a human walking the route. One of its deployments has logged more than 33,000 completed inspections across 450 separate points, a volume that only matters if nearly every one of those inspections actually finished correctly.
The Math That Turns a Good Demo Into a Bad Deployment
The arithmetic behind that 37 percent figure is not exotic, which is exactly why it is so easy for a procurement team to overlook. A single action succeeding 99 percent of the time sounds, intuitively, like a system that is basically finished. Compound that same 99 percent across 100 sequential actions in a real shift, picking, placing, navigating around an obstacle, re-grasping a dropped item, and the probability of completing the entire sequence without a single failure collapses to about 0.99 raised to the 100th power, or roughly 36.6 percent. Push per-step reliability to 99.9 percent and the full-sequence success rate climbs to about 90.5 percent. Push it to 99.99 percent, a bar vanishingly few deployed systems currently clear across every subtask in a real environment, and full-sequence reliability reaches about 99.0 percent. Each additional "nine" of reliability is not a marginal improvement; it is the difference between a robot that needs a human minder nearby at all times and one that genuinely does not.
That compounding effect is also why a vendor's single best demo clip is close to useless as a reliability signal, and why buyers who rely on it are routinely disappointed six months into a pilot. A ten-second video shows one short sequence executed once, under conditions the vendor controlled. It cannot show whether the underlying system holds that same success rate across the hundredth repetition of the day, the dim lighting at the end of a shift, or the slightly different crate geometry that a human worker handles without thinking and a robot has never seen. The honest reliability number a buyer needs is not "does it work," which almost every credible vendor can now demonstrate, but "how many consecutive real-world repetitions, across how much natural variation, before it fails," which almost none publish voluntarily.
Why Recovery Matters More Than the Headline Success Rate
If compounding reliability is the uncomfortable math, recovery is the more actionable engineering answer to it. A system that detects a slipping grip, a misaligned part or a navigation error early enough to correct it before the failure cascades into a stopped line does not need to be perfect at the first attempt; it needs to be good at noticing when the first attempt is going wrong. That reframes the entire evaluation question a buyer should be asking a vendor. Instead of "what is your success rate," the more useful question is "what fraction of your failures get caught and corrected by the system itself before a human has to intervene," because that number predicts how much supervision labor a deployment actually requires, which is usually the line item that determines whether automating a task pencils out at all.
Starship Technologies, the sidewalk-delivery robot operator, is a useful case study here precisely because its task looks deceptively simple from the outside: roll a small wheeled robot from a store to a doorstep. More than 10 million completed autonomous deliveries later, the company's actual engineering problem was never the nominal driving task; it was building a system that could detect and recover from the long tail of real-world interruptions, a car parked across a curb cut, a dog on a leash, a gate left shut, without routing every edge case to a remote human operator. A delivery count in the tens of millions is only an impressive number because the recovery layer underneath it kept the human-intervention rate low enough for the unit economics to work at that scale.
Matching the Robot to the Job Instead of Chasing Human Likeness
The industry's current fixation on general-purpose, human-shaped robots complicates this picture rather than simplifying it, because a bipedal humanoid has to solve dexterity, balance and locomotion simultaneously, compounding the reliability math across more subsystems at once than a purpose-built platform needs to. Agility Robotics, the Oregon-based maker of the Digit bipedal robot, has pursued warehouse tote-handling deployments specifically because that use case lets the company narrow Digit's dexterity requirements to a tractable subset, walking on flat warehouse floors, picking up and setting down totes of known dimensions, rather than the open-ended manipulation a general household humanoid would eventually need. That narrowing is not a limitation to apologize for. It is the correct engineering response to compounding reliability math, and it is the same instinct that led ANYbotics to build a four-legged inspection platform instead of a bipedal one for terrain that rewards stability over human-like gait, and Starship to build a wheeled delivery robot instead of a walking one for sidewalks that do not require legs at all.
The useful question for a procurement team is not which of these form factors is more advanced in some abstract sense. It is "minimum sufficient dexterity": what is the narrowest set of manipulation, locomotion and sensing capabilities that actually completes this specific economically valuable job, reliably, at a cost below what a human doing the same job costs. A robot engineered to that minimum, with a strong recovery layer built around exactly the failure modes its narrower task exposes it to, will clear the reliability bar faster and more cheaply than a general-purpose platform solving every one of those problems at once, even if the general-purpose platform looks more impressive in a trade-show video.
What Deployment Teaches That No Lab Benchmark Can
Every one of these three companies' current reliability levels is a product of real-world deployment, not laboratory benchmarking. A controlled lab environment cannot generate the specific, messy failure modes, the dropped item at a slightly wrong angle, the gauge partially obscured by condensation, the delivery route blocked by construction, that actually determine whether a system has closed enough of the gap between 99 percent and 99.99 percent to be trusted unsupervised. That is why deployment itself functions as the real training signal in this industry right now, more than any single new model release or sensor upgrade: every failure a system survives and corrects in the field becomes a data point that tightens the next version's sensing, control and recovery logic, compounding in the opposite direction from the reliability math above, each fielded unit effectively adding evidence toward the next nine.
That has a direct implication for how a buyer should weigh a vendor's deployment history against its technical specifications sheet. A platform with a smaller number of flashy capabilities but years of continuous field operation, ANYbotics' inspection runs, Starship's delivery volume, Agility's warehouse pilots, has had more opportunities to encounter and correct the long tail of real-world failure than a platform with a more capable-looking demo but little sustained field time. Specification sheets describe what a robot can do under ideal conditions. Deployment history describes what it has already proven it can survive under conditions nobody designed in advance, and the second number is the one that predicts whether a pilot converts into a renewed contract.
The Metric Set That Should Replace the Demo Reel
None of this argues against ambition in robot design; it argues for measuring the right things before committing a budget to any particular platform. The share of full workflows a system completes without human intervention, its self-corrected recovery rate, its performance across documented environmental variation rather than a single staged setting, its durability under continuous use, and its cost per successfully completed job are all harder to extract from a vendor than a demo clip, and all more predictive of whether a deployment survives its first year. A buyer who asks for those five numbers before signing a pilot contract will learn more about whether a system is ready for unsupervised work than a dozen more demo videos could show.
The three companies cited here are not the only ones running this experiment, and none of them would claim to have closed the gap between demo and deployment completely. What distinguishes a dexterous system genuinely approaching production readiness from one still rehearsing for investors is not a cleaner single-take video. It is a growing, documented record of a fielded machine doing the same unglamorous task thousands of times in a row, catching its own mistakes before they become someone else's problem, and quietly adding another nine.
This analysis synthesizes public statements and documented deployment data from the companies discussed and is for general information purposes only; it does not constitute investment, financial, legal, or professional advice.
Hero image credit: Agility Robotics.












