/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

Gemini co-lead Oriol Vinyals says Gemini 3's gains come from better pre-training and post-training, contradicting the idea that pre-training gains are falling

which we discussed in our NeurIPS '25 talk with @ilyasut and @quocleix—the team delivered a drastic jump. The delta between 2.5 and 3.0 is [image] Andrej Karpathy / @karpathy : I played with Gemini 3 yesterday via early access. ...  I had a positive early impression yesterday across personality, writing, vibe coding, humor, etc., very solid daily driver potential, clearly a tier 1 LLM, congrats to the team! Kyle Chan / @kyleichan : This has implications for China's AI industry and Nvidia chips. While Google's TPU program is exceptional in many ways, it shows that you can train a massive, state-of-the-art model outside of Nvidia's hardware & CUDA ecosystem. Kyle Chan / @kyleichan : This is the big story here. Google trained Gemini 3 Pro on Google's own TPUs. No mention of Nvidia chips. Jen Zhu / @jenzhuscott : Wait, Gemini 3 was trained on TPUs?! [image] Jason Lee / @jasondeanlee : Keep hearing from the GDM imo team @lmthang and @jj_at_brown @quocleix etc that the IMO gold methods are completely general purpose and not IMO specific that is attributed to an improvement in gemini, not some scaffolding. Then I try gemini 3 (first time using gemini since 2.5 Prinz / @deredleritt3r : After testing a few recently released models over the past few days, I have to apologize to OpenAI. I really liked GPT-5 Thinking when it was released, but thought of it as more or less “o3.1” (in Jerry Tworek's words) with drastically reduced hallucinations. But I was wrong. I @scaling01 : Gemini 3 Pro has around ~7.5T params (vibe-mathing with explanation) > the naive fit with with an R^2 of 0.8816 yields a mean estimation of 2.325 Quadrillion parameters > ummm, that's not it > let's only take sparse MoE reasoning models > this includes gpt-oss-20B and 120B, [image] @stochasticchasm : Pretraining believers we are so back Jonas Adler / @jonasaadler : Reports on the death of pre-training have indeed been greatly exaggerated. Cody Blakeney / @code_star : Pretraining is back baby [image] Jason Lee / @jasondeanlee : Benchmaxxed. No good at vibeproving Anton Tsitsulin / @graph_ : insider video of the Gemini pretraining team making a crack in the Ilya wall [image] Prinz / @deredleritt3r : “No walls in sight” for pre-training. Post-training is a “total greenfield”. 2026 is going to be a great year. Alex Tomala / @a__tomala : An important way we improved our models was to invent a time machine to learn what techniques ended up working well in the future. This way we can more efficiently use our TPUs for experiments. Super excited for the NeurIPS '25 talk!

The Information Stephanie Palazzolo

Context & Ripple Effects

Google had already broadened Gemini 2.5 Pro access beyond its initial paid tier, making model-quality improvements more immediately relevant to a wider user base. The reported jump in Gemini 3 extends that product arc from wider Gemini 2.5 Pro availability toward a stronger flagship-model cycle.

The claim also sits alongside Google’s effort to turn model capability into differentiated interfaces, including Gemini’s interactive dynamic-view output. Later coverage of Gemini 3 Pro’s visual-reasoning benchmarks reinforces why the training-method discussion matters: it bears on whether improvements can continue across multiple capabilities.

First-order effects

  • Google gains an internal technical case for continuing to invest in both pre-training and post-training rather than treating base-model scaling as exhausted.
  • Gemini users and enterprise buyers have a stronger reason to evaluate the new generation on broad day-to-day qualities, not only narrow benchmark results.

Second-order effects

  • Competing model labs face added pressure to show that their own training stacks can still produce large base-model gains, rather than relying primarily on product features or post-training refinements.
  • The report that Gemini 3 Pro was trained on Google TPUs strengthens the strategic value of Google’s vertically integrated compute path, while not by itself establishing a broader displacement of Nvidia hardware.

Third-order effects

  • If comparable gains continue, the frontier-model race may remain governed by access to training compute, data, and training expertise—not simply by incremental product-layer differentiation.
  • The key uncertainty is reproducibility: one model generation is evidence against a universal pre-training plateau, but not proof that returns will remain high across labs or future scale levels.

The trend: Frontier AI competition is increasingly testing whether integrated compute and training pipelines can keep delivering capability gains from both pre-training and post-training.

Discussion

  • @oriolvinyalsml Oriol Vinyals on x
    The secret behind Gemini 3? Simple: Improving pre-training & post-training 🤯 Pre-training: Contra the popular belief that scaling is over—which we discussed in our NeurIPS '25 talk with @ilyasut and @quocleix—the team delivered a drastic jump. The delta between 2.5 and 3.0 is [im…
  • @karpathy Andrej Karpathy on x
    I played with Gemini 3 yesterday via early access. ...  I had a positive early impression yesterday across personality, writing, vibe coding, humor, etc., very solid daily driver potential, clearly a tier 1 LLM, congrats to the team!
  • @kyleichan Kyle Chan on x
    This has implications for China's AI industry and Nvidia chips. While Google's TPU program is exceptional in many ways, it shows that you can train a massive, state-of-the-art model outside of Nvidia's hardware & CUDA ecosystem.
  • @kyleichan Kyle Chan on x
    This is the big story here. Google trained Gemini 3 Pro on Google's own TPUs. No mention of Nvidia chips.
  • @jenzhuscott Jen Zhu on x
    Wait, Gemini 3 was trained on TPUs?! [image]
  • @jasondeanlee Jason Lee on x
    Keep hearing from the GDM imo team @lmthang and @jj_at_brown @quocleix etc that the IMO gold methods are completely general purpose and not IMO specific that is attributed to an improvement in gemini, not some scaffolding. Then I try gemini 3 (first time using gemini since 2.5
  • @deredleritt3r Prinz on x
    After testing a few recently released models over the past few days, I have to apologize to OpenAI. I really liked GPT-5 Thinking when it was released, but thought of it as more or less “o3.1” (in Jerry Tworek's words) with drastically reduced hallucinations. But I was wrong. I
  • @scaling01 @scaling01 on x
    Gemini 3 Pro has around ~7.5T params (vibe-mathing with explanation) > the naive fit with with an R^2 of 0.8816 yields a mean estimation of 2.325 Quadrillion parameters > ummm, that's not it > let's only take sparse MoE reasoning models > this includes gpt-oss-20B and 120B, [imag…
  • @stochasticchasm @stochasticchasm on x
    Pretraining believers we are so back
  • @jonasaadler Jonas Adler on x
    Reports on the death of pre-training have indeed been greatly exaggerated.
  • @code_star Cody Blakeney on x
    Pretraining is back baby [image]
  • @jasondeanlee Jason Lee on x
    Benchmaxxed. No good at vibeproving
  • @graph_ Anton Tsitsulin on x
    insider video of the Gemini pretraining team making a crack in the Ilya wall [image]
  • @deredleritt3r Prinz on x
    “No walls in sight” for pre-training. Post-training is a “total greenfield”. 2026 is going to be a great year.
  • @a__tomala Alex Tomala on x
    An important way we improved our models was to invent a time machine to learn what techniques ended up working well in the future. This way we can more efficiently use our TPUs for experiments. Super excited for the NeurIPS '25 talk!