Google says Gemini 3.6 Flash improves coding, multimodal, and knowledge work performance and uses up to 17% fewer tokens and costs less per token vs. 3.5 Flash
Alphabet is releasing three new Gemini models on Tuesday, including its clearest answer yet to Anthropic's lead in cybersecurity …
CNBCMacKenzie Sigalos
Context & Ripple Effects
Google has been moving its Flash line along a recurring performance-per-cost path, from Gemini 3 Flash’s claimed faster, lower-cost reasoning to a Flash-Lite variant positioned as a cheaper option. The new release extends that positioning into coding, multimodal work and knowledge tasks.
The wider launch also includes a cyber-focused Flash model and the start of a Gemini 4 pre-training run, as reported in Google’s broader Gemini 3.6 and cyber-model rollout. That makes efficiency a central part of Google’s competitive response, not just a single-model benchmark claim.
First-order effects
Google can offer Gemini 3.6 Flash users a model it says improves key work-oriented tasks while generating up to 17% fewer output tokens and lowering per-token cost versus 3.5 Flash.
Developers and enterprise buyers using Gemini APIs gain another reason to test or migrate workloads where token consumption and model quality are both material procurement criteria.
Second-order effects
Lower token use and pricing pressure can force rival model providers, including Anthropic, to sharpen their own price-performance positioning in coding, multimodal and security-adjacent workloads.
For application builders, cheaper inference can make more capable default model settings viable, while increasing the importance of measuring total task cost rather than list price alone.
Third-order effects
If successive Flash releases keep improving capability while reducing inference cost, frontier-model competition will increasingly be decided by efficient deployment at scale rather than by benchmark leadership alone.
The pattern strengthens the tiered Gemini product strategy: specialized, lower-cost variants can segment workloads by latency, capability and security needs, potentially tightening platform lock-in for API customers.
The trend: This is one data point in the compute-to-API flywheel, where model vendors turn efficiency gains into lower-cost products and broader developer adoption.
...We believe it's one of the most powerful model available today, compatible to leading frontier AI models , second only to Fable 5. You don't have to wait to test it. Just now, the Qwen3.8-Max-Preview made its debut on Alibaba's Token Plan, Qoder, and QoderWork. [image]
Gemini 3.6 Flash is live in Antigravity! ⚡️ Building on 3.5 Flash feedback, it consumes up to 17% fewer output tokens while completing complex workflows in fewer reasoning steps and tool calls. [image]
Say hello to Gemini 3.6 Flash, designed to be higher intelligence, more token efficient, and with a new lower price, based directly on developer feedback! 3.6 Flash continues our progress towards models that are deeply usable in real world scenarios! [image]
We're rolling out three new models to make AI agents faster, smarter, and cheaper at scale: 🔵 Gemini 3.6 Flash: It uses fewer tokens than 3.5 Flash to deliver higher quality work at the exact same cost. 🔵 Gemini 3.5 Flash-Lite: A fast, cost-effective option for everyday tasks [im…
🆕 @GoogleAI's Gemini 3.6 Flash is now generally available and rolling out in GitHub Copilot. ➡️ It is designed for web and app development, coding and agentic tasks ➡️ In testing, it demonstrated higher task-completion rates and better token efficiency than Gemini 3.5 Flash
Its release day. Gemini 3.6 Flash - It's cheaper than Gemini 3.5 Flash ($7.50 output instead of $9.00 output). - It outperforms Gemini 3.1 Pro in almost every benchmark. It's being compared to GPT-5.6 Luna and Sonnet 5. Priced between these two models, it generally performs [imag…
Gemini 3.6 Flash benchmarks are out, and it's... beaten by other models on code tasks, and is only really consistently SoTA on vision and context benchmarks. But hey, 3.1 Pro is now so old 3.6 Flash outperforms it across the board 😭 [image]
⚡️Gemini 3.6 Flash is now available in Google AI Studio, and it's cheaper than Gemini 3.5 Flash. Pricing: - Input: $1.50 - Output: $7.50 (Gemini 3.5 Flash: $9.00) Knowledge cutoff: March 2026. [image]
Gemini 3.6 Flash and 3.5 Flash Lite now available in OpenCode - 1M context - 3.6 Flash: 17% cheaper output than 3.5 Flash - 3.5 Flash Lite: 80% cheaper than 3.5 Flash
Gemini 3.6 Flash is slightly more token-efficient than 3.5 but idk why you would take anything but Grok 4.5 or Sol on lower reasoning settings right now [image]
Told you all Gemini 3.6 Flash is worse than grok 4.5 in all coding task Sonnet 5 > Grok 4.5 > GPT 5.6 Luna > Gemini 3.6 Flash At this point I genuinely want google to be serious , I was one of the guy who praises google last year was so hyped now I think it's just lack of [image]
Flash has now lapped 3.5 Pro, which is still AWOL. (Worse, so has Meta?!) With the touting of ‘3.5 Flash-Cyber’ - not to mention the mention of Gemini 4 pre-train work starting, seems fair to wonder if 3.5 Pro is a dud, like Llama 4 ‘Behemoth’ before it.. spyglass.org/google-ge…