Google DeepMind says Gemini Diffusion, an experimental text diffusion model demoed at Google I/O and available by waitlist, generates 1,000-2,000 tokens/second
Our state-of-the-art, experimental text diffusion model Jose Antonio Lanz / Decrypt : Google Doubles Down on AI: Veo 3, Imagen 4 and Gemini Diffusion Push Creative Boundaries Matthias Bastian / The Decoder : Gemini Diffusion could be Google's most important I/O news that slipped under the radar Andrew Romero / 9to5Google : AI videos from Google's Veo 3, which features synchronized audio as well as crisper, more detailed video, are both impressive and terrifying Kyle Wiggers / TechCrunch : Google tests bringing ads to AI Mode, both below responses as well as “integrated into” them, and expands ads in AI Overviews to desktop in the US Bluesky: Tim Kellogg / @timkellogg.me : oh wow, Gemini is doing is doing a text diffusion model — this is likely most useful when you have a fixed peak amount of time you can wait for a response, like in robotics — blog.google/technology/g... Tim Kellogg / @timkellogg.me : I got access to Gemini Diffusion. It definitely has small model feels, but i like it — long responses appear in evenly-sized chunks. so i think they're doing like 1000 tokens at a time. i did not anticipate that but it makes sense [embedded post] Mastodon: Simon Willison / @simon@fedi.simonwillison.net : I got access to Gemini Diffusion, Google's first diffusion LLM, and the thing is absurdly fast - it ran at 857 tokens/second and built me a prototype chat interface in just a couple of seconds, video here: https://simonwillison.net/... X: Brendan O'Donoghue / @bodonoghue85 : Excited to share what my team has been working on lately - Gemini diffusion! We bring diffusion to language modeling, yielding more power and blazing speeds! 🚀🚀🚀 Gemini diffusion is especially strong at coding. In this example the model generates at 2000 tokens/sec, [video] Oriol Vinyals / @oriolvinyalsml : Today we introduced Gemini Diffusion⚡️ (& DeepThink, Veo3, Imagen4, 2.5 updates...). It's been a dream of mine to remove the need for “left to right” text generation. It's so fast, that we had to *slow down* the video during the presentation. https://deepmind.google/... [video] Jack Rae / @jack_w_rae : The Gemini Diffusion release feels like a landmark moment. For text generation, autoregressive models have always outperformed diffusion models from a quality perspective. It wasn't clear that the gap could ever be closed. The team behind this have kept laser focused, broken [image] Alexander Doria / @dorialexander : Gemini Diffusion does pass honorably my nearly impossible OCR correction benchmark. [image] @kalomaze : ughhhh how do they determine adequate depth. i dont wanna be in the sampling mines again if this approach catches on Sebastian Flennerhag / @flennerhag : Excited to share what we've been cooking - Gemini Diffusion!⚡️ Super proud of the team - cracking text diffusion was never a given but now the door is open for new capabilities and unparalleled speed. You can experience vibe-coding in real time here: https://deepmind.google/... [video] Prateek Jain / @jainprateek_ : Thrilled about Gemini Diffusion, the SOTA text diffusion model! Generates 2000 tokens/sec while outperforming Flash-lite on coding tasks. Really ambitious, innovative project with endless possibilities, and the team is just getting started! Congrats @bodonoghue85, @flennerhag, [image] @hillbig : Gemini Diffusion employs diffusion models for LLM, achieving nearly 5x faster output (1500 tokens/second). While diffusion LLMs are particularly effective for coding where fill/fix-in-the-middle are common, this will likely catch up in other domains. https://deepmind.google/... Deedy / @deedydas : The future of building software. LLMs are pretty good at generating code, but they're slow. Gemini Diffusion is 10-15x faster than autoregressive models by using diffusion, which used to be for images. This is the 2nd model after Mercury Small to show this. 2/12 Blanca Huergo / @blancahuergo : Very excited to share what I have been working on. Having been part of the Gemini Diffusion team since day one, it is amazing to see our model demoed at Google I/O :) sign up below to try it out! @testingcatalog : Gemini Diffusion will be one of the steps that will define how user interfaces will evolve within the next several years. This is a huge prerequisite for the next level of generative UIs. There won't be a need to build a frontend UIs soon. Insane 🤯 [video] Edouard Leurent / @eleurent : Excited to share what I've been up to: Gemini Diffusion is FAST! I'm convinced this will revolutionise iterative workflows: refine, get instant feedback, repeat! So proud of what our small team achieved here🪐 [video] John Lindquist / @johnlindquist : The Future of Development: Gemini Diffusion [video] Jean Tarbouriech / @jean_tarbou : 1000+ words per second! ⚡ We just unleashed Gemini Diffusion at #GoogleIO! 🚀 Awesome being part of the team that took this from a small research project all the way to I/O @GoogleDeepMind 🪐 [video] @kimmonismus : [video] Gemini Diffusion has also been lost among the announcements. However, nobody would have expected a diffusion model with 2000t/s to be on a par with Transformer models. It is particularly good at coding. Underhyped. @pminervini : Since Gemini Diffusion was just announced, diffusion LLMs may become mainstream in the near future! Being able to incorporate arbitrary constraints into the model can unlock many possibilities in terms of trustworthiness and robustness 🚀🚀🚀 Paper: https://arxiv.org/... [image] @lauriewired : I tried to tell you guys that dLLMs were cool😉 gemini diffusion is a neat coder! @googleai : Gemini Diffusion, our newest research model, is significantly faster than our fastest model so far AND matches its coding performance. By correcting errors as the model thinks, it is extremely fast for editing tasks like math and coding. [image] Wes Roth / @wesrothmoney : I just coded up 7 apps in 30 seconds with Gemini Diffusion...this has to be a world record. the video is 1x speed 👀 [video] Volodymyr Kuleshov / @volokuleshov : Congratulations Google on announcing a Mercury-level diffusion language model! 🙃 https://deepmind.google/... Petar Veličković / @petarv_93 : little known fact: i sit next to the team that built gemini diffusion — such an amazing and dedicated group of people, keeping us inspired every day! — and they've now delivered this amazing model for google i/o. give it a try... it's blazing fast! 🚀♊️🪐 Pranam Chatterjee / @pranamanam : As you know, we've been deep in the discrete diffusion trenches for sequence design for quite some time now (both theoretically and for biology!)—PepTune for multi-objective discrete diffusion for therapeutic peptide design from @_sophia_tang_, P2 for path planning and improved Cindy Wu / @cindyxywu : My team at GDM has just launched Gemini Diffusion, a SOTA text diffusion model, at #GoogleIO. Text diffusion generates the text in parallel via iterative refinement, making it super fast. Get on the waitlist to try the experimental demo model: https://deepmind.google/... @stanfordnlp : It's an interesting phenomenon of the current age how development of large deep learning text models (LLMs) is sucking in the research brainpower of so many)! Sander Dieleman / @sedielem : In 2022, I worked on text diffusion for a bit and wrote a blog post. Since then, people have regularly asked me about scaling diffusion LLMs. All the while, I was on the first row watching Brendan assemble a cracked team and make it a reality. Now I can stop being coy about it😁 Amy Lu / @amyxlu : It's finally happening!!! Diffusion is so much more satisfying than autoregressive for protein & DNA sequences that don't really have directionality 🥹 Waiting for this to empirically land & replace BERT/one-step discrete diffusion for protein foundation models 👀 Jeremy Howard / @jeremyphoward : @bodonoghue85 Makes me so happy to see this! :D I've hearing about this project for quite some time, and was really hoping that it would see the light of day. Archie Sengupta / @archiexzzz : Got access to Google diffusion. HOLY SH!T 909 tokens/s ?????? I made a calendar in 3s? 3 fcuking seconds? [image] @googledeepmind : We've developed Gemini Diffusion: our state-of-the-art text diffusion model. Instead of predicting text directly, it learns to generate outputs by refining noise, step-by-step. This helps it excel at coding and math, where it can iterate over solutions quickly. #GoogleIO [image] LinkedIn: Aaron Gokaslan : It's so exciting to see our recent work on discrete diffusion deployed into production less than a year after we published Masked Diffusion Language Models. … Lian Jye Su : Lots of announcements from #Google I/O, but the one that really got me excited was the launch of #Gemini Diffusion. … Blanca Huergo : Very excited to share what I have been working on. Having been part of the Gemini Diffusion team since day one, it is amazing to see our model demoed at Google I/O :) sign up below to try it out! … Antonio Gulli : Among all the gems released at Google I/O, I cherry pick #Gemini #Diffusion for three reasons: It's so unbelievably fast that feels unreal … Forums: Hacker News : Gemini Diffusion r/LocalLLaMA : Why nobody mentioned “Gemini Diffusion” here? It's a BIG deal r/mlscaling : Gemini Diffusion See also Mediagazer
Context & Ripple Effects
Google has progressively widened Gemini access, from the long-context developer and enterprise release of Gemini 1.5 to a broader rollout of Gemini 2.5 Pro Experimental. Gemini Diffusion shifts the emphasis from model breadth and reasoning toward response-generation speed.
That matters most for workflows built around repeated edits rather than a single answer. Its waitlist status keeps the claimed performance in an experimental phase, but the coding and math focus makes it relevant to Google's subsequent push to put Gemini closer to developers' local work.
First-order effects
- Waitlisted users can test a text model that Google says produces roughly 1,000–2,000 tokens per second, potentially reducing the visible delay in iterative coding and math edits.
- Google DeepMind gains a distinct performance claim for Gemini alongside its existing reasoning-model line, rather than positioning every task around one general-purpose model.
Second-order effects
- If the speed holds in broader use, coding-assistant and agent products will have stronger incentives to route edit-heavy tasks to fast models, making latency a more explicit product differentiator alongside output quality.
- Google's developer tooling can benefit from quicker model turns: the later Gemini CLI launch shows the company is connecting Gemini to local codebases, where repeated request-and-edit loops are central.
Third-order effects
- The result points toward AI competition being segmented by inference profile—fast iterative generation for some workloads, deeper reasoning for others—rather than settled by a single benchmark leader.
- If diffusion-style text models prove reliable at scale, serving speed and cost efficiency could become more important infrastructure advantages, while developers may increasingly select models task by task.
The trend: Frontier AI platforms are moving from one-model positioning toward specialized inference systems optimized for distinct latency, reasoning, and workflow needs.