Gemini co-lead Oriol Vinyals says Gemini 3's gains come from better pre-training and post-training, contradicting the idea that pre-training gains are falling
which we discussed in our NeurIPS '25 talk with @ilyasut and @quocleix—the team delivered a drastic jump. The delta between 2.5 and 3.0 is [image] Andrej Karpathy / @karpathy : I played with Gemini 3 yesterday via early access. ... I had a positive early impression yesterday across personality, writing, vibe coding, humor, etc., very solid daily driver potential, clearly a tier 1 LLM, congrats to the team! Kyle Chan / @kyleichan : This has implications for China's AI industry and Nvidia chips. While Google's TPU program is exceptional in many ways, it shows that you can train a massive, state-of-the-art model outside of Nvidia's hardware & CUDA ecosystem. Kyle Chan / @kyleichan : This is the big story here. Google trained Gemini 3 Pro on Google's own TPUs. No mention of Nvidia chips. Jen Zhu / @jenzhuscott : Wait, Gemini 3 was trained on TPUs?! [image] Jason Lee / @jasondeanlee : Keep hearing from the GDM imo team @lmthang and @jj_at_brown @quocleix etc that the IMO gold methods are completely general purpose and not IMO specific that is attributed to an improvement in gemini, not some scaffolding. Then I try gemini 3 (first time using gemini since 2.5 Prinz / @deredleritt3r : After testing a few recently released models over the past few days, I have to apologize to OpenAI. I really liked GPT-5 Thinking when it was released, but thought of it as more or less “o3.1” (in Jerry Tworek's words) with drastically reduced hallucinations. But I was wrong. I @scaling01 : Gemini 3 Pro has around ~7.5T params (vibe-mathing with explanation) > the naive fit with with an R^2 of 0.8816 yields a mean estimation of 2.325 Quadrillion parameters > ummm, that's not it > let's only take sparse MoE reasoning models > this includes gpt-oss-20B and 120B, [image] @stochasticchasm : Pretraining believers we are so back Jonas Adler / @jonasaadler : Reports on the death of pre-training have indeed been greatly exaggerated. Cody Blakeney / @code_star : Pretraining is back baby [image] Jason Lee / @jasondeanlee : Benchmaxxed. No good at vibeproving Anton Tsitsulin / @graph_ : insider video of the Gemini pretraining team making a crack in the Ilya wall [image] Prinz / @deredleritt3r : “No walls in sight” for pre-training. Post-training is a “total greenfield”. 2026 is going to be a great year. Alex Tomala / @a__tomala : An important way we improved our models was to invent a time machine to learn what techniques ended up working well in the future. This way we can more efficiently use our TPUs for experiments. Super excited for the NeurIPS '25 talk!
Context & Ripple Effects
Google had already broadened Gemini 2.5 Pro access beyond its initial paid tier, making model-quality improvements more immediately relevant to a wider user base. The reported jump in Gemini 3 extends that product arc from wider Gemini 2.5 Pro availability toward a stronger flagship-model cycle.
The claim also sits alongside Google’s effort to turn model capability into differentiated interfaces, including Gemini’s interactive dynamic-view output. Later coverage of Gemini 3 Pro’s visual-reasoning benchmarks reinforces why the training-method discussion matters: it bears on whether improvements can continue across multiple capabilities.
First-order effects
- Google gains an internal technical case for continuing to invest in both pre-training and post-training rather than treating base-model scaling as exhausted.
- Gemini users and enterprise buyers have a stronger reason to evaluate the new generation on broad day-to-day qualities, not only narrow benchmark results.
Second-order effects
- Competing model labs face added pressure to show that their own training stacks can still produce large base-model gains, rather than relying primarily on product features or post-training refinements.
- The report that Gemini 3 Pro was trained on Google TPUs strengthens the strategic value of Google’s vertically integrated compute path, while not by itself establishing a broader displacement of Nvidia hardware.
Third-order effects
- If comparable gains continue, the frontier-model race may remain governed by access to training compute, data, and training expertise—not simply by incremental product-layer differentiation.
- The key uncertainty is reproducibility: one model generation is evidence against a universal pre-training plateau, but not proof that returns will remain high across labs or future scale levels.
The trend: Frontier AI competition is increasingly testing whether integrated compute and training pipelines can keep delivering capability gains from both pre-training and post-training.