This Robot Learns a New Task From One Short Video, No Training Required
Generalist AI's GEN-1.5 takes a three to twelve second demonstration as a prompt and performs the task at 59 percent success, rising to 83 percent after five minutes of data. All numbers come from the company.
Robotics startup Generalist AI unveiled GEN-1.5 on 20 August, a model that picks up a new physical task from a single demonstration. You show it a clip of three to twelve seconds, the clip goes into the model’s context window as what the company calls a physical prompt, and the robot then attempts the task with no training step at all. Across ten tests, including opening a jar and pulling money out of a wallet, Generalist reports an average success rate of 59 percent. Add ten training steps on five minutes of data and that rises to 83 percent.
The context window is worth translating, because it is doing the heavy lifting here. It is simply how much information a model can hold in mind at once while it works, the machine equivalent of short-term memory. Normally you teach a robot a task by collecting hours of examples and retraining it, which takes days. GEN-1.5 instead treats the demonstration as part of the question, the same way you might paste a document into a chatbot rather than fine-tuning the model on it. Generalist says the system can chain two prompts into longer sequences, accept demonstrations recorded in simulation rather than the real world, and partly imitate human hand movements. The company also claims these abilities were never explicitly trained, but emerged during more than eight months of pretraining on interaction data.
That last claim is the one to hold lightly. The technique itself, called in-context learning, is well established in language models and other groups have shown it for robots before, though only across a handful of task types. Generalist’s claim is that it now works broadly. But the tasks demonstrated are short and simple, every number comes from the company’s own testing, and nothing has been independently verified. A 59 percent success rate also sounds better than it feels: a jar that opens six times out of ten is not a jar you would let a robot open in your kitchen. What makes this interesting is the direction rather than the score, because if showing replaces retraining, the cost of teaching a robot anything drops from days to seconds.
What this means for you: nothing today, and that is fine. There is no product here, no price and no shipping date. The reason to pay attention is that robotics has been stuck for years on the sheer cost of collecting training data for every single task, and “just show it once” is the shape of the fix everyone has been looking for. If you already work with AI tools, the useful transfer is conceptual: putting an example directly into the prompt often beats trying to retrain something, and that holds for text and images as much as for robot arms.
Sources
Google Adds a Button That Lets Readers Pick Their Own Favourite Sources
Publishers can now embed a Preferred Sources button on their own site. Google says readers are twice as likely to click through to a source they have marked as preferred.