Google DeepMind, Meta, Nvidia, and others are racing to release world models, aiming to navigate the physical world by learning from videos and robotic data
Google DeepMind, Meta and Nvidia are developing systems that aim to better understand the physical world
Context & Ripple Effects
This competition builds on Meta's decision to open-source V-JEPA 2 for predicting 3D environments and motion, moving world models from a research direction toward a platform choice for robotics and autonomous systems.
Google DeepMind had already described AutoRT methods that pair robots with visual-language models, while Nvidia's position links model development to the compute and deployment stack. The significance is that several AI leaders are now converging on the same physical-world capability layer.
First-order effects
- Google DeepMind, Meta and Nvidia face a more explicit race to demonstrate that their models can learn useful representations of physical environments from video and robot data.
- Robotics and autonomous-system developers gain multiple prospective model suppliers, rather than treating physical-world understanding as a capability developed entirely in-house.
Second-order effects
- Competition will shift toward differentiated access to training data, simulation, robotics integrations and the compute needed to train and run these models—not just benchmark performance.
- Meta's open-model posture raises pressure on proprietary offerings to show clearer deployment advantages, while hardware providers can position optimized infrastructure around physical-AI workloads.
Third-order effects
- If these efforts translate into reliable deployments, world models could become a shared foundation layer between general-purpose AI and embodied products, concentrating value in data, evaluation and integration ecosystems.
- The field may also sharpen a split between broadly available model weights and tightly controlled end-to-end stacks; the outcome depends on whether open models can match proprietary systems in real-world reliability.
The trend: World models are becoming a new competitive layer in AI, as leading labs seek to extend learned perception and prediction from digital content into physical systems.