Physical AIAI01 sources

Skild's S1 learns a ten-minute task from one human video

A black robotic gripper with an exposed camera module grasps the stem of a leafy green plant, a second robot arm blurred in the background of an indoor room, with the words Introducing S1 overlaid at left.

Skild AI / newsroomPress kit

Skild AI announced S1 on 25 August 2026, a robot foundation model prompted with video rather than language. A single egocentric human demonstration goes in as the prompt; the model executes the task with no fine-tuning and no change to its weights.

On tasks absent from pretraining, at 100,000 hours of pretraining data, Skild reports 66 percent success against 9 percent for an equivalent language-conditioned policy. On tasks inside the training distribution it reports 96 percent against 53. The company puts the exchange rate at roughly 380 post-training episodes per video demonstration. Demonstrated tasks run up to ten minutes and dozens of steps: potting a plant, making pour-over coffee, frying pancakes, assembling a kit. Skild says the model does not replay the demonstration but improvises around perturbations and its own mistakes.

What is published is a company blog post. There is no paper, no public weights and no API, only a sign-up form for early access, so none of these numbers has been reproduced outside Skild. The claim is that showing a robot a task once is now worth more than teaching it hundreds of times, which is a claim about the cost of every future deployment.

Sources

  1. [1]Introducing S1: In-Context Learning for RoboticsSkild AI··Newsroom