Physical AIAI01 sources

AgiBot trains an action model from scratch on robot data

A diagram from AgiBot's announcement showing three panels of training data — human manipulation video, robot attempts, and teleoperated demonstrations — feeding a central model, with a photograph of a humanoid robot reaching across a table on the right.

AgiBot — GE-Act 2.0 announcementPress kit

AgiBot announced GE-Act 2.0 on 9 September 2026, a model that predicts both how a scene will change and what the robot should do to change it. The company’s claim is about where it started: the model was trained from random initialisation on robot operation data, rather than taking a video-generation model trained on internet footage and adapting it to control a body.

That distinction has been the open argument in robot learning. Video models have seen enormous amounts of the world, but they have watched it rather than acted in it, and what they know about consequences is inferred from pixels. AgiBot’s staged pretraining used roughly 3.9 million hours of instruction video, 3.2 million hours of robot trajectories, and 3 million hours of matched video-and-action pairs.

The scaling figures are the substance. Raising training data from 300 hours to 30,000 hours took one platform from 39 completed tasks to 76, with success rates climbing from 17.1% to 44.1%; a second platform went from 24 tasks to 72. Evaluation covered 100 atomic tasks in 20 skill categories, tested without fine-tuning or demonstrations.

A 44% success rate is not a working robot. It is a curve that has not yet bent, which is a different kind of claim.

Sources

  1. [1]智元发布GE-Act 2.0,首次验证原生世界—动作模型预训练与Scaling路径AgiBot (智元机器人)··Press release