CourionAI
EN
Newsletter
← All news
deepmind 2 min read

Google DeepMind's new robot brain is meant to run everything from a desk arm to a humanoid

Gemini Robotics 2 and Gemini Robotics ER 2 split the job of controlling a robot into thinking and moving. Here is what a vision-language-action model actually is, and what is still on a waitlist.

A tabletop arm, a wheeled unit and a walking humanoid frame all wired up to one shared brain shape above them

Google DeepMind has introduced Gemini Robotics 2, which it calls its most advanced vision-language-action model so far, together with a second model called Gemini Robotics ER 2. The pitch is that one model can drive very different machines, from a small arm bolted to a table up to a full-body humanoid, and that it can handle whole-body movement, fine motor work, and several robots working together.

A vision-language-action model, usually shortened to VLA, is worth unpacking, because the name is the explanation. It takes in pictures from the robot’s cameras, takes in instructions in ordinary language, and outputs actions, meaning the actual motor commands that move joints and grippers. Before VLAs, those three things were usually separate systems glued together, and the glue was where most robots fell apart. Putting them in one model is what lets you say “put the mug in the sink” instead of scripting every centimetre of the motion.

The second model, Gemini Robotics ER 2, does a different job. ER stands for embodied reasoning, which is the less glamorous but arguably harder half: understanding a physical scene and deciding what should happen next. It sits above the action model as a planner, and it replaces Gemini Robotics ER 1.6 from April. ER 2 is available to try in Google AI Studio. Gemini Robotics 2 itself is not generally available, and developers have to apply through a waitlist.

What is behind this. The split into two models is the point, and it mirrors what is happening elsewhere in AI. A big, slow, expensive model does the thinking; a faster, more specialised model does the doing. Microsoft is arguing for the same division of labour in text models, and Anthropic ships Fable 5 as a manager that delegates. In robotics the split is even more necessary, because a robot arm cannot wait several seconds for a reasoning model to finish a paragraph. Worth keeping expectations grounded, though: DeepMind’s numbers come from its own demonstrations, there is no independent benchmark for “can this humanoid do laundry”, and a waitlist is a long way from a product.

What this means for you: for most people, nothing today, and that is fine. Nobody is putting a Gemini-powered humanoid in a flat this year. If you are curious, ER 2 is the piece you can actually poke at in AI Studio, and it is a good way to see how a model reasons about a physical scene. If you follow robotics professionally, the thing to watch is not the humanoid demos but the claim about controlling many different robot bodies with one model. Robotics has always been bespoke, one control stack per machine, and a general layer that ports across hardware would change the economics far more than any single impressive video.

Sources

Source: https://deepmind.google/models/gemini-robotics/

Next story

Mira Murati's lab shrank its own model to a quarter of the size and lost one point

Thinking Machines released Inkling Small, an open-weights reasoning model with 276 billion parameters that scores 40 where the 975-billion-parameter original scores 41.

A balance scale holding one huge hollow cube against one small dense cube, resting almost level