Google says Gemini 3.6 Flash improves coding, multimodal, and knowledge work performance and uses up to 17% fewer tokens and costs less per token vs. 3.5 Flash
Google has steadily positioned its Flash line around making capable multimodal models cheaper and faster to deploy, from the earlier lightweight Gemini 1.5 Flash to Gemini 3 Flash's claimed move toward higher-end reasoning at lower latency.
The latest release extends that arc by pairing capability claims with lower API pricing; Google separately listed lower Gemini 3.6 Flash token prices than the prior Flash version, making efficiency a product-level differentiator rather than just a benchmark result.
First-order effects
Developers using Gemini 3.5 Flash have a lower-cost successor to evaluate for coding, multimodal, and knowledge-work workloads, with Google claiming up to 17% lower token use as well as lower per-token pricing.
Google strengthens the commercial position of its Flash tier by making the performance-per-token proposition central to the upgrade.
Second-order effects
Competing model providers face more pressure to show both task quality and effective inference cost, since customers can compare API bills alongside headline model performance.
Lower token consumption can reduce the cost of applications with repeated or long-running model interactions, potentially widening the workloads for which developers consider a faster, lower-priced model tier viable.
Third-order effects
If vendors continue improving capability while cutting token use and API rates, model selection will increasingly turn on effective inference cost—output quality adjusted for the tokens and price required to achieve it—rather than raw model labels alone.
That dynamic could favor providers able to translate infrastructure and model-efficiency gains into frequent, credible price-performance upgrades, though actual adoption will depend on independent workload results.
The trend: This is one data point in the AI API market's shift from selling larger models on peak capability toward competing on efficient, production-ready performance per dollar.
We're rolling out three new models to make AI agents faster, smarter, and cheaper at scale: 🔵 Gemini 3.6 Flash: It uses fewer tokens than 3.5 Flash to deliver higher quality work at the exact same cost. 🔵 Gemini 3.5 Flash-Lite: A fast, cost-effective option for everyday tasks [im…
Its release day. Gemini 3.6 Flash - It's cheaper than Gemini 3.5 Flash ($7.50 output instead of $9.00 output). - It outperforms Gemini 3.1 Pro in almost every benchmark. It's being compared to GPT-5.6 Luna and Sonnet 5. Priced between these two models, it generally performs b…
🆕 @GoogleAI's Gemini 3.6 Flash is now generally available and rolling out in GitHub Copilot. ➡️ It is designed for web and app development, coding and agentic tasks ➡️ In testing, it demonstrated higher task-completion rates and better token efficiency than Gemini 3.5 Flash
Gemini 3.6 Flash and 3.5 Flash Lite now available in OpenCode - 1M context - 3.6 Flash: 17% cheaper output than 3.5 Flash - 3.5 Flash Lite: 80% cheaper than 3.5 Flash
Say hello to Gemini 3.6 Flash, designed to be higher intelligence, more token efficient, and with a new lower price, based directly on developer feedback! 3.6 Flash continues our progress towards models that are deeply usable in real world scenarios! [image]
Told you all Gemini 3.6 Flash is worse than grok 4.5 in all coding task Sonnet 5 > Grok 4.5 > GPT 5.6 Luna > Gemini 3.6 Flash At this point I genuinely want google to be serious , I was one of the guy who praises google last year was so hyped now I think it's just lack of [image]
Gemini 3.6 Flash is slightly more token-efficient than 3.5 but idk why you would take anything but Grok 4.5 or Sol on lower reasoning settings right now [image]
⚡️Gemini 3.6 Flash is now available in Google AI Studio, and it's cheaper than Gemini 3.5 Flash. Pricing: - Input: $1.50 - Output: $7.50 (Gemini 3.5 Flash: $9.00) Knowledge cutoff: March 2026. [image]
Flash has now lapped 3.5 Pro, which is still AWOL. (Worse, so has Meta?!) With the touting of ‘3.5 Flash-Cyber’ - not to mention the mention of Gemini 4 pre-train work starting, seems fair to wonder if 3.5 Pro is a dud, like Llama 4 ‘Behemoth’ before it.. spyglass.org/google-ge…