/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

Google says Gemini 3.6 Flash improves coding, multimodal, and knowledge work performance and uses up to 17% fewer tokens and costs less per token vs. 3.5 Flash

MacKenzie Sigalos /CNBC:

CNBC MacKenzie Sigalos

Context & Ripple Effects

Google has steadily positioned its Flash line around making capable multimodal models cheaper and faster to deploy, from the earlier lightweight Gemini 1.5 Flash to Gemini 3 Flash's claimed move toward higher-end reasoning at lower latency.

The latest release extends that arc by pairing capability claims with lower API pricing; Google separately listed lower Gemini 3.6 Flash token prices than the prior Flash version, making efficiency a product-level differentiator rather than just a benchmark result.

First-order effects

  • Developers using Gemini 3.5 Flash have a lower-cost successor to evaluate for coding, multimodal, and knowledge-work workloads, with Google claiming up to 17% lower token use as well as lower per-token pricing.
  • Google strengthens the commercial position of its Flash tier by making the performance-per-token proposition central to the upgrade.

Second-order effects

  • Competing model providers face more pressure to show both task quality and effective inference cost, since customers can compare API bills alongside headline model performance.
  • Lower token consumption can reduce the cost of applications with repeated or long-running model interactions, potentially widening the workloads for which developers consider a faster, lower-priced model tier viable.

Third-order effects

  • If vendors continue improving capability while cutting token use and API rates, model selection will increasingly turn on effective inference cost—output quality adjusted for the tokens and price required to achieve it—rather than raw model labels alone.
  • That dynamic could favor providers able to translate infrastructure and model-efficiency gains into frequent, credible price-performance upgrades, though actual adoption will depend on independent workload results.

The trend: This is one data point in the AI API market's shift from selling larger models on peak capability toward competing on efficient, production-ready performance per dollar.

Discussion

  • @googledeepmind @googledeepmind on x
    We're rolling out three new models to make AI agents faster, smarter, and cheaper at scale: 🔵 Gemini 3.6 Flash: It uses fewer tokens than 3.5 Flash to deliver higher quality work at the exact same cost. 🔵 Gemini 3.5 Flash-Lite: A fast, cost-effective option for everyday tasks [im…
  • @kimmonismus @kimmonismus on x
    Its release day.  Gemini 3.6 Flash - It's cheaper than Gemini 3.5 Flash ($7.50 output instead of $9.00 output).  - It outperforms Gemini 3.1 Pro in almost every benchmark.  It's being compared to GPT-5.6 Luna and Sonnet 5.  Priced between these two models, it generally performs b…
  • @adamholtererer Adam Holter on x
    Gemini 3.6 Flash Frontend Test: Without Skills vs. With Skills [image]
  • @github @github on x
    🆕 @GoogleAI's Gemini 3.6 Flash is now generally available and rolling out in GitHub Copilot. ➡️ It is designed for web and app development, coding and agentic tasks ➡️ In testing, it demonstrated higher task-completion rates and better token efficiency than Gemini 3.5 Flash
  • @opencode @opencode on x
    Gemini 3.6 Flash and 3.5 Flash Lite now available in OpenCode - 1M context - 3.6 Flash: 17% cheaper output than 3.5 Flash - 3.5 Flash Lite: 80% cheaper than 3.5 Flash
  • @officiallogank Logan Kilpatrick on x
    Say hello to Gemini 3.6 Flash, designed to be higher intelligence, more token efficient, and with a new lower price, based directly on developer feedback! 3.6 Flash continues our progress towards models that are deeply usable in real world scenarios! [image]
  • @adamholtererer Adam Holter on x
    Gemini 3.6 Flash Benchmarks: [image]
  • @chetaslua @chetaslua on x
    Told you all Gemini 3.6 Flash is worse than grok 4.5 in all coding task Sonnet 5 > Grok 4.5 > GPT 5.6 Luna > Gemini 3.6 Flash At this point I genuinely want google to be serious , I was one of the guy who praises google last year was so hyped now I think it's just lack of [image]
  • @scaling01 @scaling01 on x
    Gemini 3.6 Flash is slightly more token-efficient than 3.5 but idk why you would take anything but Grok 4.5 or Sol on lower reasoning settings right now [image]
  • @pankajkumar_dev Pankaj Kumar on x
    Gemini 3.6 Flash is now available in Google AI Studio. Pricing: $1.50/M input • $7.50/M output Knowledge cutoff: March 2026 [image]
  • @ai_for_success AshutoshShrivastava on x
    ⚡️Gemini 3.6 Flash is now available in Google AI Studio, and it's cheaper than Gemini 3.5 Flash. Pricing: - Input: $1.50 - Output: $7.50 (Gemini 3.5 Flash: $9.00) Knowledge cutoff: March 2026. [image]
  • @synthwavedd Leo on x
    DeepMind should just surrender all of their compute to Moonshot wtf is this 💔
  • @mgsiegler.com M.G. Siegler on bluesky
    Flash has now lapped 3.5 Pro, which is still AWOL.  (Worse, so has Meta?!)  With the touting of ‘3.5 Flash-Cyber’ - not to mention the mention of Gemini 4 pre-train work starting, seems fair to wonder if 3.5 Pro is a dud, like Llama 4 ‘Behemoth’ before it.. spyglass.org/google-ge…