World Labs Shows Atlas, a Model That Builds Walkable 3D Scenes From a Handful of Photos
Fei-Fei Li's company unveiled a model that generates, reconstructs, and simulates 3D space rather than flat video. It outputs point clouds, runs a minute of 1440p footage, and can produce training data for robots.
World Labs, the company co-founded by Fei-Fei Li, announced a model last week called Atlas that does something most AI image and video tools cannot: it understands where things are in space. Feed it two or three photos of a room and it can show you that room from angles nobody photographed, output it as actual 3D data, and let a simulated robot walk through it.
The technical difference is small to describe and large in effect. Video models treat a scene as a sequence of flat frames. Atlas anchors every input, whether text, image, video, or 3D data, to a position in three-dimensional space first. World Labs calls this “spatial context.” It means camera movement is passed in as geometry, an actual path you draw, rather than described in a prompt and hoped for.
The concrete numbers: Atlas outputs up to one minute of video at 1440p. It reconstructs scenes from as few as one image and handles more than a hundred. In one demo it assembles Stanford’s Main Quad from 25 ground-level photos and then generates aerial views from far above. It can export point clouds and Gaussian splats, which are ways of storing a scene as many small points in space so it can be viewed smoothly from any direction. In head-to-head tests judged by human evaluators, reviewers preferred Atlas over MiniMax H3 in 75 percent of comparisons and over Seedance 2.5 in 94 percent.
What is behind this
Li has been arguing for a while that “spatial intelligence” is the missing piece. Language models flatten the world into a line of tokens, video models into a stack of frames, and both then struggle with questions a toddler finds easy, like what is behind that chair. Atlas is the attempt to build the missing layer directly.
The most commercially interesting use is not pretty pictures. It is robots. Atlas can reconstruct a real room and then generate what a robot’s cameras and depth sensors would see along any path through it, with objects, lighting, and backgrounds swapped around. That turns one real recording into thousands of training situations, which is the expensive bottleneck in robotics right now.
Worth keeping expectations grounded: World Labs Atlas is in early access for selected partners, there is no public benchmark that captures what it does, and most of the comparisons come from World Labs itself.
What this means for you: Nothing today, unless you work in robotics, architecture, games, or film. But it is a good marker of where the field is heading. The last three years were about models that read and write. The next stretch is about models that know where things are, which is what any machine has to understand before it can safely move through your kitchen.
Sources
Anthropic Has Signed Up to $517 Billion of Computing Power in Eleven Months
An analysis by The Information puts Anthropic's compute commitments since October at 14.8 gigawatts and as much as $517 billion over the next decade, up from the roughly $180 billion it had told investors to expect.