/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

Google says Gemini 3.6 Flash improves coding, multimodal, and knowledge work performance and uses up to 17% fewer tokens and costs less per token vs. 3.5 Flash

Alphabet is releasing three new Gemini models on Tuesday, including its clearest answer yet to Anthropic's lead in cybersecurity …

CNBC MacKenzie Sigalos

Context & Ripple Effects

Google has been moving its Flash line along a recurring performance-per-cost path, from Gemini 3 Flash’s claimed faster, lower-cost reasoning to a Flash-Lite variant positioned as a cheaper option. The new release extends that positioning into coding, multimodal work and knowledge tasks.

The wider launch also includes a cyber-focused Flash model and the start of a Gemini 4 pre-training run, as reported in Google’s broader Gemini 3.6 and cyber-model rollout. That makes efficiency a central part of Google’s competitive response, not just a single-model benchmark claim.

First-order effects

  • Google can offer Gemini 3.6 Flash users a model it says improves key work-oriented tasks while generating up to 17% fewer output tokens and lowering per-token cost versus 3.5 Flash.
  • Developers and enterprise buyers using Gemini APIs gain another reason to test or migrate workloads where token consumption and model quality are both material procurement criteria.

Second-order effects

  • Lower token use and pricing pressure can force rival model providers, including Anthropic, to sharpen their own price-performance positioning in coding, multimodal and security-adjacent workloads.
  • For application builders, cheaper inference can make more capable default model settings viable, while increasing the importance of measuring total task cost rather than list price alone.

Third-order effects

  • If successive Flash releases keep improving capability while reducing inference cost, frontier-model competition will increasingly be decided by efficient deployment at scale rather than by benchmark leadership alone.
  • The pattern strengthens the tiered Gemini product strategy: specialized, lower-cost variants can segment workloads by latency, capability and security needs, potentially tightening platform lock-in for API customers.

The trend: This is one data point in the compute-to-API flywheel, where model vendors turn efficiency gains into lower-cost products and broader developer adoption.

Discussion

  • @alibaba_qwen @alibaba_qwen on x
    ...We believe it's one of the most powerful model available today, compatible to leading frontier AI models , second only to Fable 5.  You don't have to wait to test it.  Just now, the Qwen3.8-Max-Preview made its debut on Alibaba's Token Plan, Qoder, and QoderWork.  [image]
  • @antigravity @antigravity on x
    Gemini 3.6 Flash is live in Antigravity! ⚡️ Building on 3.5 Flash feedback, it consumes up to 17% fewer output tokens while completing complex workflows in fewer reasoning steps and tool calls. [image]
  • @officiallogank Logan Kilpatrick on x
    Say hello to Gemini 3.6 Flash, designed to be higher intelligence, more token efficient, and with a new lower price, based directly on developer feedback! 3.6 Flash continues our progress towards models that are deeply usable in real world scenarios! [image]
  • @googledeepmind @googledeepmind on x
    We're rolling out three new models to make AI agents faster, smarter, and cheaper at scale: 🔵 Gemini 3.6 Flash: It uses fewer tokens than 3.5 Flash to deliver higher quality work at the exact same cost. 🔵 Gemini 3.5 Flash-Lite: A fast, cost-effective option for everyday tasks [im…
  • @github @github on x
    🆕 @GoogleAI's Gemini 3.6 Flash is now generally available and rolling out in GitHub Copilot. ➡️ It is designed for web and app development, coding and agentic tasks ➡️ In testing, it demonstrated higher task-completion rates and better token efficiency than Gemini 3.5 Flash
  • @kimmonismus @kimmonismus on x
    Its release day. Gemini 3.6 Flash - It's cheaper than Gemini 3.5 Flash ($7.50 output instead of $9.00 output). - It outperforms Gemini 3.1 Pro in almost every benchmark. It's being compared to GPT-5.6 Luna and Sonnet 5. Priced between these two models, it generally performs [imag…
  • @adamholtererer Adam Holter on x
    Gemini 3.6 Flash Benchmarks: [image]
  • @adamholtererer Adam Holter on x
    Gemini 3.6 Flash Frontend Test: Without Skills vs. With Skills [image]
  • @synthwavedd Leo on x
    Gemini 3.6 Flash benchmarks are out, and it's... beaten by other models on code tasks, and is only really consistently SoTA on vision and context benchmarks. But hey, 3.1 Pro is now so old 3.6 Flash outperforms it across the board 😭 [image]
  • @pankajkumar_dev Pankaj Kumar on x
    Gemini 3.6 Flash is now available in Google AI Studio. Pricing: $1.50/M input • $7.50/M output Knowledge cutoff: March 2026 [image]
  • @ai_for_success AshutoshShrivastava on x
    ⚡️Gemini 3.6 Flash is now available in Google AI Studio, and it's cheaper than Gemini 3.5 Flash. Pricing: - Input: $1.50 - Output: $7.50 (Gemini 3.5 Flash: $9.00) Knowledge cutoff: March 2026. [image]
  • @opencode @opencode on x
    Gemini 3.6 Flash and 3.5 Flash Lite now available in OpenCode - 1M context - 3.6 Flash: 17% cheaper output than 3.5 Flash - 3.5 Flash Lite: 80% cheaper than 3.5 Flash
  • @scaling01 @scaling01 on x
    Gemini 3.6 Flash is slightly more token-efficient than 3.5 but idk why you would take anything but Grok 4.5 or Sol on lower reasoning settings right now [image]
  • @chetaslua @chetaslua on x
    Told you all Gemini 3.6 Flash is worse than grok 4.5 in all coding task Sonnet 5 > Grok 4.5 > GPT 5.6 Luna > Gemini 3.6 Flash At this point I genuinely want google to be serious , I was one of the guy who praises google last year was so hyped now I think it's just lack of [image]
  • @synthwavedd Leo on x
    DeepMind should just surrender all of their compute to Moonshot wtf is this 💔
  • @mgsiegler.com M.G. Siegler on bluesky
    Flash has now lapped 3.5 Pro, which is still AWOL.  (Worse, so has Meta?!)  With the touting of ‘3.5 Flash-Cyber’ - not to mention the mention of Gemini 4 pre-train work starting, seems fair to wonder if 3.5 Pro is a dud, like Llama 4 ‘Behemoth’ before it.. spyglass.org/google-ge…