The Pulse
Figure says Helix 2.5 worked across 30 unfamiliar homes
Checked the Figure article for an attributable quotation from a named person and found none.

AI.info Team ·
Figure says its Helix 2.5 humanoid system has performed three household tasks across 30 Bay Area homes it had never seen, without collecting new data or adapting the model in those locations. The company describes the September 17, 2026 result as its first large-scale demonstration of whole-body humanoid generalization in unfamiliar homes.
The evaluation covered living-room tidying, towel folding, and bed making. Figure says Helix 2.5 used one foundation model pretrained on Index, the company’s dataset of human behavior, then adapted that model to the three tasks using data collected outside the test homes. The robot had to work with each home’s existing furniture and with objects that did not appear in the task-specification data.
Thirty homes, three tasks, no local retraining
Figure defines “zero-shot” narrowly in the experiment: the robot had not seen the evaluation environments or the objects it manipulated during training, and no rollout data from those homes was used to choose or update the checkpoint. Each task used a single fixed model across all 30 homes.
Success also required completing the full task rather than receiving credit for partial progress. For living-room tidying, the robot had to pick up all 13 to 15 toys placed in the scene and put them in a basket. Towel folding required every towel to be folded and placed in the basket. Bed making required both pillows and the comforter corners to reach the top of the bed, with the comforter pulled smooth.
Those tasks force the machine to combine walking, visual perception, grasping, two-handed manipulation and body positioning. A robot cannot treat the work as a fixed tabletop routine: it must locate objects, move through constrained rooms, adjust its stance and recover when an action goes wrong.
Index pretraining drives the reported jump
Figure ran an ablation to separate the effect of Index pretraining from the task-specific training data. The company trained two policies with identical downstream data, architecture, optimization settings and evaluation procedures. One began with random weights; the other began with the Index-pretrained Helix 2.5 model.
In Figure’s blind evaluations, the policy trained from scratch succeeded on 9% of zero-shot trials. The Index-pretrained policy succeeded on 56%, a result the company describes as more than six times higher. Figure says no single evaluation task represented more than 1.90% of the Index pretraining dataset.
The result supports Figure’s claim that broad human-behavior data, rather than task demonstrations alone, supplies much of the model’s ability to transfer a learned behavior into a new physical setting. The company does not present the 56% figure as a solution to household robotics; its own account says general humanoid robotics is not solved.
Half the adaptation data for a wider test
Figure also compared Helix 2.5 with a previous Helix 02 policy trained for the same task. Helix 02 relied on data collected directly in the environment where it was evaluated. Figure says Helix 2.5 matched that policy’s success rate while using half as much adaptation data, then operated across 30 homes it had not seen.
The company highlights self-correction as a qualitative difference. In its description, the robot steps back to reposition itself, changes stance and moves around a bed when a fold needs correction. Those behaviors matter because long household tasks accumulate small errors, particularly when furniture and object placement vary from one home to another.
A scaling claim tied to human video
Figure trained four models on nested subsets of Index covering an eightfold increase in pretraining data, while holding model size and downstream training constant. Each model was fine-tuned on the same task data and evaluated using the same held-out next-action prediction loss.
The company reports that loss declined predictably with each doubling of Index data. Using the smaller training runs, Figure says it forecast the largest run’s test loss to four decimal places before running that experiment. Forecasting error amounted to 0.54% of the variation across the full eightfold data range.
That experiment measures data scaling only. It does not establish how changes in model size, compute or downstream training would affect performance. Figure presents the result as evidence that human-to-robot transfer may be predictable enough to guide future data collection and training decisions.
Figure’s next bottleneck is scale
Figure says Index now generates roughly 35 minutes of new human experience every second and that the company has committed $3.5 billion in compute to training Helix. The stated strategy is to learn more physical behavior before a robot enters a new home, reducing the need to collect demonstrations in every deployment environment.
Helix 2.5’s reported 56% success rate still leaves a large share of strict all-or-nothing trials incomplete. The result matters because the test moves beyond a controlled room: a single checkpoint must handle unfamiliar layouts, furniture and objects while coordinating the robot’s full body. For now, Figure’s concrete claim is narrower than a household-ready machine: three specified behaviors, 30 unseen homes and no local adaptation.