CourionAI
EN
Newsletter
← All news
robotics 2 min read

This Robot Learns a Ten-Minute Task From Watching One Video

Skild AI released S1, a robot model that copies a task after seeing a single human demonstration, with no retraining. It flipped a pancake despite never seeing pancake flipping in training.

A robot arm flipping a pancake in a pan, with a filmstrip beside it showing the same motion frame by frame

Skild AI released a robotics model called S1 on August 25 with an unusual claim: show it one video of a person doing something, and the robot does that thing, even if the task takes ten minutes and even if nothing like it appeared in the model’s training. No retraining, no data collection session, no engineer in the loop.

The demonstrations Skild published cover pancake flipping, pour-over coffee, potting a plant and assembling a kit, all multi-step tasks with a lot of small decisions in the middle. The company says that when S1 flipped a pancake for the first time, the team went back through the training data and found no examples of flipping at all, meaning the model worked it out from the single video prompt. On tasks it had never seen, Skild reports a 66 percent success rate against 9 percent for language-prompted VLA models, short for vision-language-action models, the current standard approach where you describe a task in words instead of showing it. Trained at the same scale of roughly 100,000 hours of robot data. Skild’s own framing of the gap: to reach what S1 does from one video, a conventional model would first need 50 to 100 hours of collected demonstrations and then a fine-tuning run, meaning a round of extra training on that specific task.

What’s actually going on here: this is robotics borrowing a trick that language models learned years ago. When you paste an example into a chatbot and it copies the pattern for the rest of the conversation, that is in-context learning: the model adapts from what is in front of it without any of its internal settings changing. Robotics has mostly not worked that way. Each new task meant new data and a new training run, which is precisely why useful robots stayed stuck in factories doing one thing forever. Swapping the prompt from a sentence to a video also sidesteps a stubborn problem, which is that language is a terrible way to specify physical motion. “Flip the pancake” leaves out everything that matters. A five-second clip does not.

What this means for you: nothing in your kitchen changes this year, and it is worth noting these are company-published demos rather than independent testing, with a success rate that still means one attempt in three fails. But the direction is the point. If teaching a robot a new job stops being an engineering project and becomes a matter of filming yourself doing it once, the economics of the whole field change, and so does the list of jobs that involve hands. If you work anywhere near warehousing, manufacturing, catering or care, this is the research thread worth following over the next two years, not the humanoid robot videos that get shared for the sci-fi look of them.

Sources

Source: https://www.skild.ai/blogs/s1

Next story

Alibaba's Next Open Model Is a Preview of Qwen 4

Qwen3.8-Flash-Next is a 125-billion-parameter open-weight model that only uses six billion of them per word. It lands today, and it is built on the architecture behind the coming Qwen 4 family.

A wide honeycomb grid where only a handful of cells are filled in, with an hourglass at the edge