Figure releases Helix 2.5 and tests it in 30 unseen homes
▶ Figure — YouTubeVideo frame
Figure released Helix 2.5 on 17 September 2026 and reported running it in 30 Bay Area homes where the company had collected no data. Robots driven by the model tidied living rooms, folded towels and made beds. The toys, towels and bedding used in the evaluation did not appear in the task-specification data either.
Figure reports a 56% zero-shot success rate for a policy pretrained on Index — the crowdsourced archive of human chore footage it has been paying contributors to film — against 9% for an equivalent policy trained from scratch. It also says Helix 2.5 matched the success rate of an earlier Helix 02 policy while using half as much task-specific adaptation data.
The figures come from Figure and from nobody else. There is no technical report, no third-party replication, and no published per-task breakdown, so a single 56% stands in for three chores of visibly different difficulty — folding a towel and making a bed are not the same problem. What the announcement does establish is the shape of the test: the robot was put into houses whose contents it had not seen, told what to do, and got it right slightly more often than not.
