Google prices Gemini 3.6 Flash lower than 3.5 Flash, at $1.50/1M input tokens and $7.50/1M output tokens, and Gemini 3.5 Flash-Lite at $0.30/1M and $2.50/1M
As we wait for 3.5 Pro, Google today announced Gemini 3.6 Flash and 3.5 Flash-Lite, while providing updates on what comes next.
Context & Ripple Effects
Google’s Flash line has repeatedly paired smaller or faster models with lower prices, including an earlier Flash-8B variant positioned with a 50% price cut and higher rate limits.
The immediate backdrop is Gemini 3.5 Flash’s $1.50 input and $9 output pricing in May; the higher 3.5 Flash output rate made the new output-price reduction the material change in this release.
First-order effects
- Google lowers Gemini 3.6 Flash’s output-token price to $7.50 per million from Gemini 3.5 Flash’s $9, while holding the input rate at $1.50 per million.
- Gemini 3.5 Flash-Lite establishes a separate low-cost tier at $0.30 per million input tokens and $2.50 per million output tokens, giving API buyers a cheaper option for cost-sensitive workloads.
Second-order effects
- Developers using output-heavy applications can reduce Gemini spend by moving from 3.5 Flash to 3.6 Flash, subject to their own quality and migration testing.
- The widened Flash/Flash-Lite menu makes model routing more attractive: customers can reserve the higher-priced tier for tasks that need it and direct routine requests to Flash-Lite.
Third-order effects
- If successive Flash releases keep improving price-performance, API-model competition will increasingly turn on effective inference cost rather than a single published token rate.
- A more segmented portfolio can make Google’s model platform stickier, but it also raises the importance of clear workload benchmarks and routing tools for customers comparing tiers.
The trend: This is another step in the shift toward tiered, continually repriced AI inference offerings built around matching workloads to cost and latency targets.