Google DeepMind unveils Genie 2, a model that can generate 3D worlds from a single prompt image, playable by humans or AI agents using keyboard and mouse inputs
Google's Genie 2 can turn text into a playable game in real-time William J. Broad / New York Times : Google DeepMind unveils GenCast, an AI weather model that the company claims outperforms traditional methods on up to 15-day weather and deadly storm forecasts X: Jack Parker-Holder / @jparkerholder : Introducing 🧞Genie 2 🧞 - our most capable large-scale foundation world model, which can generate a diverse array of consistent worlds, playable for up to a minute. We believe Genie 2 could unlock the next wave of capabilities for embodied agents 🧠. [video] Demis Hassabis / @demishassabis : @elonmusk Thanks Elon! let's do an AI game together... Demis Hassabis / @demishassabis : The world model is taking shape... 🌐 @googledeepmind : Introducing Genie 2: our AI model that can create an endless variety of playable 3D worlds - all from a single image. 🖼️ These types of large-scale foundation world models could enable future agents to be trained and evaluated in an endless number of virtual environments. → [video] Alberto Rosas / @albertorosasg : The metaverse is closer than ever Emad / @emostaque : Every pixel will be generated, the energy cost for this is moderate & will be minimal by the time the next console upgrade cycle hits Jeff Clune / @jeffclune : In my 2019 AI-GA paper I proposed a neural net world model as a “Darwin Complete” environment search space that could produce any possible environment for open-ended learning. It felt like a flight of fancy. I knew rationally it was possible eventually, but emotionally it felt Nick Dobos / @nickadobos : remember like a month ago when haters laughed at the ai minecraft demo? “this is useless why would anyone play ai minecraft, whats the point??” this: Alex Volkov / @altryne : Don't believe we live in a simulation yet anon? “Genie 2 is a world model, meaning it can simulate virtual worlds, including the consequences of taking any action (e.g. jump, swim, etc.). It was trained on a large-scale video dataset and, like other generative models, Philippe Lemoine / @phl43 : I have long been thinking that training models on textual data was a serious limitation, and that we'd eventually need to rely on richer sensory-type data to overcome them, but it hadn't occurred to me that we could do that by using embodied agents in a virtual world. @junglesilicon : when you close your eyes and get told to imagine a beach, you see it, hear it, smell it. for just a moment it'll feel like you're there. text is good at expressing some things but isn't the right medium for others. we simulate the world in our mind cos it's useful. @reach_vb : DeepMind COOKED! Genie 2, a large-scale, multi-modal foundation world model! 🔥 Capable of creating endless action-controllable, playable 3D environments - the future is going to be so, so wild! [video] Alex Tabarrok / @atabarrok : Probability that we are living in a simulation just went up. Robert Scoble / @scobleizer : Yet another Holodeck. The 3D innovations are falling out of the sky. Oh, literally true on my desk. The sky is above my X Pro and Macintosh monitor. These posts fall right out of the sky onto my monitor. Heh. And I'm not even wearing my Apple Vision Pro. When I do life gets Nikshep Saravanan / @nikshepsvn : big upgrade from the original genie, going from playable 2D platformers to actual 3D RPGs genie: https://sites.google.com/... genie 2: https://deepmind.google/... Bogdan Mazoure / @bogdan_mazoure : Great progress towards general interactive world models! Another reason why you should leverage advances in large generative models for decision making Yixin Lin / @yixin_lin_ : These state-of-the-art world models clearly encode intuitive physics/3D understanding, and reason about the effects of actions. They also point to a clear path to scaling beyond games to the whole Internet. I'm confident these methods have a role for real-world robotic control. Shane Gu / @shaneguml : In 2022, I paused all my research in RL / robotics in simulation, and moved back to large GenAI because of text-to-video/3d (Imagen Video and DreamFusion). I predicted actionable text-to-video models will play a bigger role. Congratulations to Jack and the team! Ashley Edwards / @ashrewards : Wow this is amazing!! Congrats to the team on such awesome results. Moving the marker again 🎉 Jakob Foerster / @j_foerst : Genius (and looks fun, too)! This is a good example of a team (or teams) working together under an ambitious blue sky vision. Academia and industry need more of it. Riz Virk / @rizstanford : AI generation of limitless worlds is moving faster than anyone expected ... simulation point is approaching fast! Julian Togelius / @togelius : Tim's team did it again! This looks like a big step forward, and I'm excited to dig into the details. Will be very interesting to compare this to the @worldlabs demo from the other day, in terms of e.g. expressivity and consistency. Oliver Groth / @omgroth : Amazing results from the Genie project. 🤯 Something which always amazed me about the work @GoogleDeepMind, especially as a big video game fan, is how advances in ‘gaming technology’ can unlock so many other impactful applications. Work and play can go so nicely together here. 😊 Christos Kaplanis / @ckaplanis1 : It's been a real joy working with such an energised and collaborative team to build Genie 2, with which we've made a significant step towards building a general, action-controllable world model for training embodied AGI. I can't wait to see what the future holds in this space. Jessica Yung / @jessicayung17 : Excited to present Genie 2, a world model that can generate diverse, action-controllable worlds from single images. 🕹️✨ We even had the SIMA agent follow language instructions in our generated environments - a step towards training agents in unlimited, varied settings! Yuge Shi / @yugeten : Excited to finally share Genie 2! This year we unlock: ✅ any image -> 3D games ✅ keyboard & mouse control ✅ fast video generation in 720p ✅ RL agent deployment Such a great ride this past year ticking these off one after another. So proud of what we've accomplished! Ryan Sullivan / @ryansullyvan : It's incredible how big of a leap in quality this is from the first iteration of Genie! I'm excited to see what kind of curricula we could design with a fully promptable environment generator. Richard Seroter / @rseroter : What the what. The possibilities for this seem massive. Michael Dennis / @michaeld1729 : It's been a crazy 2 years seeing so many amazingly talented researchers bring GENerative Interactive Environments alive in Genie 1 and 2. the future is agents in generative environments Sundar Pichai / @sundarpichai : Incredible progress ahead! Alexandre Moufarek / @amoufarek : Excited to introduce Genie 2: our most capable and general foundation world model. It can generate a diverse array of consistent and interactive 3D worlds from a single image, playable with keyboard and mouse inputs for up to a minute.🧞🎮🌎 [video] Jeff Clune / @jeffclune : Thrilled to share Genie 2! Endless environments that can be created by text or images, a key to open-ended and AI-Generating Algorithms. Genie 1 showed it's possible. 9 months later, Genie 2 shows jaw-dropping progress.🤯 Witness the magic of scale, again. 📈🚀 It enables Tim Rocktäschel / @_rockt : Excited to reveal Genie 2, our most capable foundation world model that, given a single prompt image, can generate an endless variety of action-controllable, playable 3D worlds. Fantastic cross-team effort by the Open-Endedness Team and many other teams at @GoogleDeepMind! 🧞 [video] Bonnie Li / @bonniesjli : Super excited to share our breakthrough foundation world model!🤯 Genie 2 can generate interactive 3D worlds given any image. Amazing work with @jparkerholder and @_rockt 🚀 Elon Musk / @elonmusk : @demishassabis Cool Elon Musk / @elonmusk : @demishassabis Ok, that would be cool Aaron Pitters / @aaronpitters : This reminds me of what @Sama said back in April. We've seen updates from a few companies regarding similar tools. It had inspired me earlier in the year to reenvision what was possible with a #TVseries I was pitching. The following is from the Bible for the show. [image] Andrew Curran / @andrewcurran_ : This is DeepMind's SIMA following instructions within Genie 2's generative world. This is the same training direction Dr. Fan's team is pursuing at NVIDIA for their embodied systems. Most robotic training won't happen in base reality. [video] Shalev Lifshitz / @shalev_lif : @jparkerholder Really amazing work! We can finally combine open-ended agents and open-ended world models. We're getting closer to the near-infinite training data regime. LinkedIn: Alex Rutter : Introducing Genie 2: our AI model that can create an endless variety of playable 3D worlds - all from a single image. 🖼️ … Jake Bruce : Excited to share our latest work on Genie, a neural network generative model of interactive environments. … Forums: Hacker News : Genie 2: A large-scale foundation world model r/gamedev : Whats everyones take on Deepminds Genie 2? r/Games : Genie 2: A large-scale foundation world model ("games" generated from a single input image) r/StableDiffusion : Genie 2: A large-scale foundation world model r/singularity : Genie 2: A large-scale foundation world model
Context & Ripple Effects
Genie 2 arrives as the interactive extension of image-to-3D work: World Labs’ early image-generated game-like scenes appeared days earlier, while DeepMind positions its system as a foundation model for environments rather than a conventional game-production tool.
The later Genie 3 release with longer continuous interaction makes Genie 2 a meaningful waypoint in DeepMind’s effort to turn generated visuals into persistent, navigable worlds.
First-order effects
- Google DeepMind can use Genie 2 to generate short, consistent 3D environments from a single image, with keyboard-and-mouse interaction available to both people and AI agents.
- Developers and researchers gain a new way to prototype interactive environments and test agents without first building each world through conventional asset pipelines.
Second-order effects
- Companies pursuing image-to-3D and generative simulation, including the segment previewed by World Labs, face a higher bar: outputs must support interaction and consistency, not merely produce a convincing static scene.
- The model links generative media more closely to agent development, increasing the value of systems that can maintain an environment as an agent acts within it.
Third-order effects
- If world models extend interaction time and reliability, synthetic environments could become a core layer of embodied-agent training and evaluation rather than a niche content-generation feature.
- That shift would move competition from individual images or clips toward controllable simulations, with compute requirements and access to integrated model platforms becoming more consequential.
The trend: Generative AI is progressing from producing media assets toward producing interactive, agent-usable simulated environments.