RobotAIGeek

Figure's Helix 2.5 Cleared 30 Homes It Had Never Seen

Figure AI's Helix 2.5 model raised zero-shot task success in 30 unfamiliar homes from 9 percent to 56 percent, testing what generalization really means for home robots.

martti
4 min readPosted: Sep 18, 2026
Figure's Helix 2.5 Cleared 30 Homes It Had Never Seen

The hardest problem in home robotics has never been getting a humanoid to fold a towel once. It has been getting that same robot to fold a towel in a house it has never entered, on a bed it has never seen, using hands that have never touched that particular fabric. Figure AI says its latest control model, Helix 2.5, cleared that bar this month, running three household tasks across 30 unfamiliar homes in the San Francisco Bay Area with no additional data collection, fine-tuning, or on-site adaptation, and posting a task success rate of 56 percent, up from a 9 percent baseline for the same tasks without the model's large-scale pretraining step. For an industry that has spent the past two years proving humanoid robots can perform impressively in a single, carefully rehearsed environment, a genuine jump in zero-shot generalization across unrelated physical spaces is a different kind of result, and it is the one that determines whether home robots are a demo category or a product category.

Figure AI, the Bay Area humanoid robotics company that builds both the Figure 03 robot hardware and the Helix vision-language-action model that controls it, rented 30 homes it had never operated in before and tested three long-horizon behaviors inside each one: tidying a living room, folding towels, and making a bed, all using whatever furniture, surfaces, and objects happened to already be in that specific house. No home was used for training data, and no task-specific adaptation was performed after the robot arrived. That test design is the point. A robotics demo staged in a company's own showroom, with lighting, furniture, and object placement tuned in advance, proves a model can perform under favorable conditions. A test run cold across 30 homes the model has never seen proves something closer to what an actual customer's living room will look like on day one.

What Changed Between Helix and Helix 2.5

The mechanism behind the jump is what Figure calls Index pretraining, a large-scale pretraining stage that ingests human physical experience data at a rate the company describes as roughly 35 minutes of human experience data collected per second, well beyond what task-specific robot demonstration data alone could ever supply. Rather than training a model directly on curated, task-specific robot demonstrations for each new behavior, Figure first pretrains a single base model on that broader human-experience corpus, then adapts the same pretrained base to distinct downstream behaviors spanning locomotion, rigid and deformable object manipulation, bimanual coordination, and active perception. Holding the task-specific data, robot hardware, training procedure, and evaluation protocol fixed, and changing only whether that Index pretraining step was included, the company reports the zero-shot success rate on the 30-home evaluation rose from 9 percent to 56 percent. That is an unusually clean experimental design for a robotics announcement: most humanoid companies report a single headline number without isolating which architectural change produced it, and Figure's decision to publish the controlled comparison is itself useful evidence, because it lets outside researchers evaluate whether the pretraining approach, rather than some other undisclosed change, is actually doing the work.

The efficiency gain compounds the generalization result. Figure says Helix 2.5 needed only half as much task-specific data as a representative behavior from the earlier Helix 02 model to reach comparable or better performance, which matters because task-specific robot demonstration data, physical trials recorded with a real robot performing a real task, is the most expensive input in the entire training pipeline. Collecting that data requires robot hardware time, human operators, and often manual labeling, and it does not scale the way internet-sourced text or image data does for large language models. If pretraining on broader human-experience data genuinely substitutes for a meaningful share of that expensive task-specific collection, it changes the unit economics of training a home-robot model, not just the model's eventual performance ceiling.

The Scaling-Law Claim Underneath the Headline Number

Figure has also published a scaling-law analysis alongside the generalization result, reporting that its forecasting error for the largest pretraining run tested was just 0.54 percent of total variation across an eightfold range of pretraining data volume. In plain terms, the company is claiming it can predict, with high precision, how much a given increase in pretraining data will improve downstream task performance, before running the full-scale experiment. That kind of predictive scaling relationship is the same tool that let large language model developers justify multi-hundred-million-dollar training runs years before those models shipped, because a reliable scaling law turns an expensive training run from a speculative bet into a forecastable investment. If Figure's scaling law holds at even larger data volumes than the eightfold range tested so far, it gives the company, and its investors, a defensible basis for continuing to scale Index pretraining rather than guessing at diminishing returns.

That claim deserves the same scrutiny any single-company scaling-law announcement deserves before it is treated as settled science. Scaling laws published by a single lab, evaluated on that lab's own benchmark and hardware, have in other domains sometimes proven less robust once independent researchers tried to replicate them under different conditions, and Figure has not indicated whether Helix 2.5's scaling curve has been evaluated by outside researchers or reproduced on a different robot platform. The 30-home test, by contrast, is a harder result to argue with regardless of how the underlying scaling law eventually holds up, because 56 percent success across genuinely novel environments is either true or it is not, in a way that a projected scaling curve is not.

Where a 56 Percent Success Rate Actually Sits

A 56 percent task success rate is not a finished consumer product, and Figure has not presented it as one. Folding towels wrong, or leaving a bed half-made, in nearly half of attempts across unfamiliar homes would be a frustrating experience for an actual paying customer, and the gap between a research benchmark run once across 30 rented houses and a product that has to perform reliably every day in one specific house, indefinitely, without a camera crew or engineering team present, is substantial. What the number is useful for is comparison against the field's actual starting point rather than against an idealized target: a 9 percent baseline without Index pretraining on the exact same tasks, hardware, and evaluation protocol means the honest comparison is not "56 percent versus perfect," it is "56 percent versus the roughly one-in-ten success rate the same robot achieved a generation earlier." Judged against that baseline, a sixfold improvement in a single model iteration is a meaningfully fast rate of progress for a field that has historically improved in smaller increments between major hardware or software revisions.

The three tasks chosen for the evaluation, tidying a living room, folding towels, and making a bed, are also not the easiest possible household chores to demonstrate. They require sustained multi-step planning rather than a single grasp-and-place action, tolerance for deformable materials like towels and bedding that do not hold a fixed shape the way a rigid object does, and navigation and manipulation coordinated across a full-body, bimanual robot rather than a stationary arm. Choosing long-horizon, deformable-object tasks for a generalization benchmark, rather than a simpler pick-and-place demonstration that would likely show a higher raw success percentage, suggests Figure is optimizing its public benchmark for credibility with technically literate buyers and competitors rather than for the most flattering possible headline number, though the company obviously has every incentive to frame its own results favorably regardless of task choice.

What This Means for Everyone Chasing the Same Problem

Every company building a general-purpose humanoid for homes, from well-funded rivals with their own vision-language-action models to component suppliers building the actuators and sensors those models will eventually run on, is racing toward the same underlying constraint: task-specific training data does not scale fast enough to cover the near-infinite variety of real homes, objects, and edge cases a commercial home robot will eventually encounter. If large-scale pretraining on broader human-experience data genuinely substitutes for a meaningful share of that expensive, slow-to-collect task-specific data, as Figure's controlled 9-to-56-percent comparison suggests it might, then the competitive advantage in humanoid robotics shifts further toward whichever companies can source, process, and train on the largest volume of general physical-experience data, not just whichever company has the most capable robot hardware or the largest fleet already deployed collecting task-specific demonstrations.

That shift has a direct implication for buyers and investors evaluating humanoid robotics companies over the next several product cycles: a company's current fleet size and current task-specific dataset, the metrics most humanoid startups lead with in their own marketing, may matter less to long-term competitiveness than whether that company has, or can build, a pretraining pipeline and data-sourcing strategy comparable to Figure's Index approach. A rival with a larger deployed fleet today but no equivalent pretraining architecture could find itself generalizing far more slowly across new environments than a competitor with fewer robots in the field but a better foundation model underneath them.

The next test for Helix 2.5 is not another 30-home research evaluation. It is whichever home a paying customer eventually lets the robot into unsupervised, on an ordinary Tuesday, with no engineering team standing by and no benchmark protocol defining success in advance, doing the loads of laundry and the unmade beds that were never rented, staged, or chosen for a press release.

This analysis draws on public statements and technical results published by Figure AI regarding the Helix 2.5 model and its evaluation. It is for general information purposes only and does not constitute investment, financial, or legal advice.

Hero image credit: Figure AI.

RoboticsHumanoidRobotsPhysicalAIUSAMachineLearningRobotDeployment