Q&A with Google Gemini co-leads Jeff Dean and Noam Shazeer on Google's path to AGI, the future of Moore's Law, TPUs, inference scaling, open research, and more
“as we scale up [training], there may be a push to have a bit more asynchrony in our systems than we do now” 👀 Haider / @slow_developer : Google Chief Scientist, Jeff Dean “AI now generates 25% of Google's integrated code” Google has already trained a Gemini model on its internal codebase to help developers this doesn't cover everything, but it improves AI-assisted coding by integrating code into its parameters [video] Dwarkesh Patel / @dwarkesh_sp : The @JeffDean & @NoamShazeer episode. We talk about 25 years at Google, from PageRank to MapReduce to the Transformer to MoEs to AlphaChip - and soon to ASI. My favorite part was Jeff's vision for AGI as one giant MoE that is grown in bits and pieces over time like a forest, [video] @kimmonismus : [video] Noam Shazeer is the co-author of “Attention is All You Need”, by 2030 we have: - Personal AI assistants everywhere + Robots - World GDP to go way up to orders of magnitude higher than it is today. - We will have solved unlimited energy. Petar Veličković / @petarv_93 : From my experience, @JeffDean and @NoamShazeer are both not only stellar researchers, but also incredibly kind people that are extremely open-minded about new ideas. Well worth a listen! Dwarkesh Patel / @dwarkesh_sp : Paperback? $1 per 10,000 tokens Customer support? $1 per 100 tokens SWE/doctor/lawyer? $1 per token But LLMs give you a million tokens to the dollar (and falling) - @NoamShazeer [image] Jeff Dean / @jeffdean : Wherein I express some thoughts on making ML models attend to lots more information than they do today. Dwarkesh Patel / @dwarkesh_sp : What would it look like to combine Google search (shallow Knowledge Graph reasoning over an ultra-ultra-wide index) with LLM in-context learning (highly intelligent operations on a tiny index)? [video] @modestproposal1 : This was so good. To me these two came across as like pragmatic system engineers more than “feel the AGI” idealists. Dwarkesh keeps almost trying to bait them into the implications of their tech and they're like ok next design challenge. https://www.dwarkeshpatel.com/ ... LinkedIn: Michael David Francois : Going to be 14 years soon, and I think it is a long time, and then I am reminded of folks like Jeff Dean & Noam Shazeer, or Urs. …
Context & Ripple Effects
Google’s AI effort has moved from the Bard-and-Gemini launch period covered in an earlier discussion of Gemini’s rollout toward using its models inside its own engineering workflow. The company now says AI produces 25% of integrated code and has trained a Gemini model on its internal codebase.
The interview also ties model progress to Google’s systems strategy: TPUs, inference scaling, and a possible move toward more asynchronous training. That makes Gemini not only a user-facing model program but an internal test bed for deploying AI across a large software organization.
First-order effects
- Google’s developers gain a more codebase-aware assistant, while the reported 25% AI share makes review, integration, and quality controls more central to its software-development process.
- Google’s AI infrastructure teams must support both larger training runs and the inference workload created when coding assistance is embedded in day-to-day development.
Second-order effects
- Rival AI platforms and cloud providers face pressure to offer coding tools that can work with customers’ proprietary repositories, rather than only generate code from public patterns.
- Greater use of internal coding assistants raises the value of reliable inference capacity and hardware-software integration, reinforcing the importance of Google’s later discussion of TPUs and coding agents.
Third-order effects
- If large companies can safely tune or ground models on their own code, AI-assisted development may shift from standalone developer tools to an enterprise software-production layer with governance and repository access as key differentiators.
- The emphasis on scaling training, inference, and system asynchrony points to AI progress becoming increasingly constrained by operational compute efficiency, not model design alone.
The trend: This is part of AI industrialization: frontier-model builders are turning their own software estates and compute stacks into feedback loops for deploying and improving AI agents.