Gemini 3 hands-on: a fundamental improvement on daily use, extremely fast, Antigravity IDE is a powerful launch product, and its personality is terse and direct
Gemini 3 is a fundamental improvement on daily use, not just on benchmarks. It feels more consistent and less “spiky” than previous models.
matt shumer
Context & Ripple Effects
Gemini 3’s launch was framed around stronger coding, reasoning and factual-accuracy claims, while Google also highlighted higher reported LMArena performance and reasoning results. This hands-on account matters because it tests whether those claims translate into steadier day-to-day use rather than isolated benchmark wins.
The contrast is meaningful against an earlier Gemini Advanced hands-on assessment that found GPT-4-class capability without an obvious benchmark lead. The reported speed, consistency and more direct style suggest Google is trying to make model quality felt in routine interaction, while Antigravity IDE gives that capability a focused work surface.
First-order effects
For Gemini users, the reported reduction in erratic responses and faster interaction can make the model more practical for repeated daily tasks, not just occasional high-stakes prompts.
Antigravity IDE immediately becomes a prominent proof point for Gemini 3: its usefulness will be judged by whether the model’s coding and reasoning gains hold up inside a dedicated development environment.
Second-order effects
Competitors face pressure to compete on responsiveness, consistency and product workflow—not solely leaderboard results—as users gain a clearer experiential basis for comparing frontier models.
A strong IDE launch can pull more coding activity toward an integrated AI workspace, raising the value of tools that combine model access with task-specific interfaces.
Third-order effects
If model releases increasingly differentiate on reliability in everyday use, benchmark leadership may become less decisive than sustained performance within work surfaces such as IDEs and assistants.
This points toward AI workspace consolidation: model providers may seek to capture more of the interface where knowledge work and software development occur, though adoption will determine whether dedicated surfaces outperform general chat.
The trend: Frontier AI competition is shifting from headline benchmark gains toward faster, more dependable models embedded in purpose-built work surfaces.
this post has me lost. almost all of the advantage Google has is through DATA!!! and hardware sovereignty via TPUs, not better algos/architecture i'm sure their work there is solid, even great (esp wrt multimodal), but not “best in the world by a long shot” tier great
I've had access to Gemini 3 since November 13th. Since then, I've used it as my daily-driver, pushing it to its limits. Here's my review of Gemini 3: https://shumer.dev/...
The secret behind Gemini 3? Simple: Improving pre-training & post-training 🤯 Pre-training: Contra the popular belief that scaling is over—which we discussed in our NeurIPS '25 talk with @ilyasut and @quocleix—the team delivered a drastic jump. The delta between 2.5 and 3.0 is [im…
I played with Gemini 3 yesterday via early access. ... I had a positive early impression yesterday across personality, writing, vibe coding, humor, etc., very solid daily driver potential, clearly a tier 1 LLM, congrats to the team!
This has implications for China's AI industry and Nvidia chips. While Google's TPU program is exceptional in many ways, it shows that you can train a massive, state-of-the-art model outside of Nvidia's hardware & CUDA ecosystem.
Keep hearing from the GDM imo team @lmthang and @jj_at_brown @quocleix etc that the IMO gold methods are completely general purpose and not IMO specific that is attributed to an improvement in gemini, not some scaffolding. Then I try gemini 3 (first time using gemini since 2.5
After testing a few recently released models over the past few days, I have to apologize to OpenAI. I really liked GPT-5 Thinking when it was released, but thought of it as more or less “o3.1” (in Jerry Tworek's words) with drastically reduced hallucinations. But I was wrong. I
Gemini 3 Pro has around ~7.5T params (vibe-mathing with explanation) > the naive fit with with an R^2 of 0.8816 yields a mean estimation of 2.325 Quadrillion parameters > ummm, that's not it > let's only take sparse MoE reasoning models > this includes gpt-oss-20B and 120B, [imag…
An important way we improved our models was to invent a time machine to learn what techniques ended up working well in the future. This way we can more efficiently use our TPUs for experiments. Super excited for the NeurIPS '25 talk!