Fei-Fei Li's World Labs unveils Atlas, a multimodal world model that generates image/video frames with pixel-perfect camera control and reconstructs them in 3D
World models generate, reconstruct, and simulate any possible world. They understand how worlds appear, behave …
World Labs
Context & Ripple Effects
World Labs first previewed image-to-3D scene generation in 2024, then turned that direction into Marble's editable 3D environments in 2025. Atlas extends the arc from creating a scene to reconstructing it and controlling generated views across images and video.
The company’s $1 billion raise for world models was framed around robotics and scientific discovery. World Labs says Atlas can reconstruct spaces from a few photos and generate RGB and depth views along chosen robot trajectories, tying its visual-generation work more directly to simulation.
First-order effects
World Labs gains a single model offering for controlled image and video generation plus explicit 3D reconstruction, rather than positioning those as separate workflows.
Robotics teams evaluating synthetic sensor data can test Atlas against real-space reconstructions and selected camera trajectories, according to World Labs’ product claims.
Second-order effects
Tencent’s open-weights 3D-consistent video model becomes a more direct alternative for developers weighing model access against Atlas’s combined reconstruction and camera-control capabilities.
Tool vendors serving 3D content and robot simulation face pressure to support outputs that combine rendered views, spatial structure, and depth data rather than treating video generation as a standalone asset pipeline.
Third-order effects
If controlled world models prove useful across content creation and simulation, competition will shift from visually plausible generation toward dependable spatial representations that downstream tools can navigate and manipulate.
The category is converging on a physical-AI control plane, where camera paths, scene geometry, and sensor outputs become model inputs and outputs rather than manually assembled production artifacts.
The trend: World-model development is moving from single-image scene synthesis toward multimodal systems designed to represent, reconstruct, and simulate navigable spaces.
I'm so excited that our @theworldlabs team has achieved a major milestone today! Introducing Atlas - a first of its kind multimodal world model trained from scratch! 🚀 Atlas is capable of generating frames with pixel-perfect camera control, reconstructing large scenes from as few…
For robotic simulation, Atlas reconstructs a space from just a few photos and generates the photorealistic RGB and depth data any robot's sensors would observe on any trajectory. Robots can now be trained and tested in far more spaces. Until now, scanning spaces like these requir…
Atlas understands both the spatial structure of the world and how it evolves over time. This lets us turn a handful of ordinary cameras into a “bullet time” multiview capture studio, capturing dynamic moments frozen in time.
Atlas also outputs explicit 3D from one to many input images, beating top open source reconstruction models. Passing more images gives Atlas more context: the more it sees, the less it imagines.
By positioning multiple input views within the model's spatial context, Atlas allows you to generate image and video frames with precise control. Here, we hand-designed a camera trajectory to generate a 1 minute video at 1440p resolution from seven reference images.
We pre-trained Atlas from scratch to take multimodal inputs, including camera movement, and turn it into 3D grounded views. Atlas puts you in control: direct the views, reconstruct real spaces with your inputs, and build explorable worlds. Read more: https://www.worldlabs.ai/...
Introducing Atlas: The world's first multimodal world model that generates image and video frames with pixel-perfect camera control and reconstructs them in 3D. Model the world, move the camera, and simulate space & time.
The scaling economics of world models are different: everywhere else, verification is the bottleneck. Here you just simulate your way out of it with more compute. https://x.com/...
Atlas is an autoregressive diffusion model built from the ground up for the task of “next frame prediction”. It is simultaneously a world class method for camera-controlled video generation, novel view synthesis, and sparse 3D reconstruction.
🔥An incredible accomplishment. Think of it as a video model with full camera control. And the scene remains (nearly) 3D consistent. Built on a fully internal base model. There are many use cases, from video editing, to 3D reconstruction to robotics. Great technical blog too.
Today we sharing Atlas, our new multimodal world model. One reason this model is so special to me is that it combines two core visual intelligence tasks I have worked on for over a decade: generation and reconstruction