Google says Gemini 3 Pro scores 1,501 on LMArena, above 2.5 Pro, and demonstrates PhD-level reasoning with top scores on Humanity's Last Exam and GPQA Diamond
Google today announced Gemini 3 with the goal of bringing “any idea to life.” The first model available in this family …
9to5GoogleAbner Li
Context & Ripple Effects
Google had already used benchmark gains to frame the Gemini 2.5 Pro cycle, including an LMArena score increase for its upgraded preview. Gemini 3 Pro extends that performance-led positioning with a larger claimed lead over the prior Pro model and stronger reasoning-benchmark results.
Google gains a new, comparable performance claim for Gemini 3 Pro: 1,501 on LMArena, above Gemini 2.5 Pro, alongside its stated top results on Humanity’s Last Exam and GPQA Diamond.
Gemini’s product and developer positioning can now center on advanced reasoning as well as coding, while the reported results give users a basis to distinguish the Pro tier from Google’s earlier model.
Second-order effects
Competing model providers face added pressure to answer on shared public benchmarks and to show that reasoning gains translate into useful product behavior, rather than relying on broad capability claims.
Google can use the flagship’s reported results to support a tiered portfolio: high-end Pro performance establishes the ceiling, while lower-cost variants can be distributed more broadly, as in the later Flash default rollout.
Third-order effects
If benchmark leadership continues to be paired with broad Google distribution, frontier-model competition will increasingly turn on both measured capability and the ability to place model variants into everyday products.
The growing weight placed on arena and expert-question benchmarks may make independent evaluation and real-world reliability more important differentiators, since vendor-reported scores alone do not settle product usefulness.
The trend: This is one data point in the shift from isolated model launches toward continuously benchmarked model families whose flagship gains are leveraged across mass-distribution products.
🚀 We just launched Gemini 3 Pro — the strongest multimodal understanding model ever built. I lead product for Gemini's multimodal vision capabilities, and I want to share more about the massive wins we are seeing across document, screen, spatial, and video understanding. 🧵
Looks like x'mas came early: Gemini 3 is here! i don't put much weight on benchmarks anymore (prefer to vibe test them for my tasks), but... it looks v good beats GPT-5.1 on almost everything.
Yeah, AI progress is totally, definitely stalling... Look at MathArena Apex. GPT-5.1 scored 1%. Gemini 3 scored 23%. That is a >20x jump on one of the hardest reasoning tasks we have. But sure, keep your head in the sand... [image]
re: Gemini 3.0 Benchmarks do NOT matter and 99% of them are awful and built by idiots. There is a trillion dollar marketing war being waged for your mind and time. Everyone is paid for. Nothing matters until it does and nothing is real until you can use it.
🚨BREAKING: @GoogleDeepMind's Gemini-3-Pro is now #1 across all major Arena leaderboards 🥇#1 in Text, Vision, and WebDev - surpassing Grok-4.1, Claude-4.5, and GPT-5 🥇#1 in Coding, Math, Creative Writing, Long Queries, and nearly all occupational leaderboards. Massive gains [image…
We just verified Gemini 3 Pro and Deep Think (Preview) are over 2X SOTA on ARC v2! This is really impressive and frankly a bit surprising. Impressive because many of the v2 solves indicate clear complexity scaling over v1. Such as tasks 65b59efc, e3721c99, and dd6b8c4b We're
At Box, we've been testing Gemini 3 Pro in early access with Box AI on our most complex advanced reasoning eval, and Gemini Pro was a massive 22 percentage point improvement over Gemini 2.5 Pro. For this test, we ask the model a series of complex, real-world questions with a set …
For the first time, Google has the most intelligent model, with Gemini 3 Pro Preview improving on the previous most intelligent model, OpenAI's GPT-5.1 (high), by 3 points [image]
As well as a substantial improvement in intelligence, Gemini 3 Pro Preview has demonstrates increased token efficiency compared to Gemini 2.5 Pro, using fewer tokens on the Artificial Analysis Intelligence Index than its predecessor, as well as other leading models such as Kimi […
Gemini 3 Pro Preview leads two of the three coding evaluations in the Artificial Analysis Intelligence Index, including an impressive 56% in SciCode, an improvement of over 10 percentage points from the previous highest score. It is also strong in agentic contexts, achieving the …
Gemini 3 Pro Preview takes the top spot on the Artificial Analysis Omniscience Index, our new benchmark for measuring knowledge and hallucination across domains. Gemini 3 Pro Preview comes in first for both Omniscience Index (our lead metric that takes off points for incorrect [i…
Individual results across the 10 evals we run independently for the Artificial Analysis Intelligence Index: MMLU-Pro, GPQA Diamond, Humanity's Last Exam, LiveCodeBench, SciCode, AIME 2025, IFBench, AA-LCR, Terminal-Bench Hard, 𝜏²-Bench Telecom [image]
Its so over for OpenAI and Anthropic. Gemini 3 Pro Benchmarks 37.5% on HLE 31.1% on ARC-AGI-2 2439 Elon on LiveCodeBench Pro 85.4% on Tau-Bench 72.1% on SimpleQA Verified SOTA everywhere except SWE-Bench Verified [image]
Context Arena Update: Added Gemini 3.0 Pro Preview (Thinking, 11-18) to the MRCR leaderboards. It establishes a new state-of-the-art in context performance, taking the #1 spot on all our AUC leaderboards and for nearly all pointwise scores. All results at: [image]
Gemini Pro 3 benchmark result I'm most excited about, would love to never have to click through random booking UIs ever again (e.g. reservations, flights, etc) [image]
Gemini 3 benchmarks leaked...wow Highest in every benchmark except coding (just behind Sonnet 4.5) Truly mind blowing If your boss doesn't give you the afternoon off to play with super intelligence, you 100% should quit your job now [image]
🎉 Huge congrats to @GoogleDeepMind on the release of Gemini-3.0-Pro, which scored an impressive 72.7 on our ScreenSpot-Pro benchmark, outperforming specialized GUI-focused models by over 6 points. 🚀 Also excited to see ScreenSpot-Pro now widely adopted by almost all leading [imag…
Gemini 3.0 Pro absolutely dominates every benchmark! The jump from 2.5 is nuts! Its scores on the most difficult benchmarks suggest this is essentially baby AGI! Humanity Last exam: 37.5% ARC-AGI-2: 31.1% LiveCodeBench Pro: 2439 Math arena apex : 23.4% Simple QA: 72.1% [image]
Introducing Gemini 3 Pro, the world's most intelligent model that can help you being anything to life. It is state of the art across most benchmarks, but really comes to life across our products (AI Studio, the Gemini API, Gemini App, etc) 🤯 [image]
Introducing Gemini 3 Pro > best model in world for multimodal understanding > SOTA at complex reasoning > insanely good at agentic tool use and vibe-coding > great at zero-shot code generation > beats Claude sonnet 4.5, GPT-5.1 on almost every benchmark [image]
Gemini 3 Pro is the new leader in AI. Google has the leading language model for the first time, with Gemini 3 Pro debuting +3 points above GPT-5.1 in our Artificial Analysis Intelligence Index @GoogleDeepMind gave us pre-release access to Gemini 3 Pro Preview. The model [image]
You can look at the benchmarks and Gemini 3 is an incremental advance (like GPT-5), but it is a significant incremental advance across benchmarks #AI [embedded post]