AgiBot trains an action model from scratch on robot data

AgiBot — GE-Act 2.0 announcementPress kit
AgiBot announced GE-Act 2.0 on 9 September 2026, a model that predicts both how a scene will change and what the robot should do to change it. The company’s claim is about where it started: the model was trained from random initialisation on robot operation data, rather than taking a video-generation model trained on internet footage and adapting it to control a body.
That distinction has been the open argument in robot learning. Video models have seen enormous amounts of the world, but they have watched it rather than acted in it, and what they know about consequences is inferred from pixels. AgiBot’s staged pretraining used roughly 3.9 million hours of instruction video, 3.2 million hours of robot trajectories, and 3 million hours of matched video-and-action pairs.
The scaling figures are the substance. Raising training data from 300 hours to 30,000 hours took one platform from 39 completed tasks to 76, with success rates climbing from 17.1% to 44.1%; a second platform went from 24 tasks to 72. Evaluation covered 100 atomic tasks in 20 skill categories, tested without fine-tuning or demonstrations.
A 44% success rate is not a working robot. It is a curve that has not yet bent, which is a different kind of claim.