Google launches Gemini 3.1 Flash-Lite, which it says delivers “enhanced performance” at a fraction of the cost of larger models and outperforms 2.5 Flash
Get best-in-class intelligence for your highest-volume workloads. … Today, we're introducing Gemini 3.1 Flash-Lite …
The new model extends that product ladder by positioning a Lite tier for high-volume work while claiming performance beyond Gemini 2.5 Flash. It follows Google’s Gemini 3 Flash launch, which similarly emphasized stronger reasoning with lower latency and cost.
First-order effects
Google adds a newer low-cost Gemini option for customers whose workloads prioritize volume, giving them a potential alternative to larger models and older Flash tiers.
The performance claim puts Gemini 2.5 Flash under immediate internal product pressure: customers evaluating upgrades have a clearer reason to benchmark it against 3.1 Flash-Lite.
Second-order effects
Google’s model menu becomes more explicitly tiered, requiring customers to trade off capability, latency and unit cost across Flash, Flash-Lite and larger Gemini offerings rather than defaulting to one general-purpose model.
Competing API providers face added pressure to pair model improvements with lower-cost serving options for high-throughput applications, not solely to release flagship models.
Third-order effects
If performance gains continue to reach lower-priced tiers, AI application economics could shift toward broader deployment of inference-heavy features, with model selection increasingly governed by workload-specific cost discipline.
This reinforces a compute-to-API flywheel in which infrastructure efficiency becomes a product differentiator: providers that can turn serving gains into cheaper, capable APIs may gain more usage and feedback.
The trend: Frontier AI vendors are increasingly competing by pushing usable model capability down into cheaper, faster inference tiers for production-scale workloads.
Gemini 3.1 Flash-Lite is available now! It takes an unbelievable amount of complex engineering to make AI feel instantaneous, enabling exciting new frontiers for experimentation! [video]
smol but incredibly mighty! Gemini 3.1 Flash-Lite is an absolute speed demon (417 tokens/s!! 🏃♀️💨) but still punches far above its weight class. 💸 $0.25 per 1M input tokens 🧠 86.9% on GPQA Diamond ⚡️ 2.5X faster time-to-first-token A game changer for high-volume and agent
Here's a side-by-side speed and accuracy comparison comparing Gemini 3.1 Flash Lite against our older Gemini 2.5 Flash model. Not only is the new model significantly faster in terms of tokens/s, it also needs many fewer tokens to accomplish complex tasks (~1/3 as many in this [vi…
Today, we're introducing Gemini 3.1 Flash Lite (in preview) ⚡️ Now available via the Gemini API, our fastest and most cost-efficient Gemini 3 series model: - Features dynamic thinking for scaled reasoning - Delivers enhanced performance at a lower cost (priced at $0.25/1M input […
Gemini 3.1 Flash-Lite Just Dropped & Destroyed the Budget Model Tier Google's cheapest model and it's somehow the fastest in its class. The numbers at $0.25 input: > 363 tok/s — Flash-Lite > 108 tok/s — Claude Haiku ($1.00) > 71 tok/s — GPT-5 mini ($0.25) Adjustable “thinking [vi…
⚡ Excited to announce Gemini 3.1 Flash-Lite! We've set a new standard for efficiency and capability to give developers our fastest, most cost-effective Gemini 3 model yet. We engineered this model with thinking levels, allowing it to handle high-volume queries instantly, while [i…
Smarter. Faster. Gemini 3.1 Flash-Lite is here⚡ The model offers uncompromising speed & intelligence at scale by focusing on: — Cost-efficiency: Priced at just $0.25/1M input and $1.50/1M output tokens, it gets work done faster at a fraction of the cost of larger models, [video]
Gemini 3.1 Flash-Lite is here! Our fastest, most cost-efficient gemini model built for high-volume workloads at scale. 💰 Priced at $0.25/M input, $1.50/M output tokens 🧠 Matches 2.5 Flash quality at Flash-Lite cost ⚡2.5x TFT and 45% faster output vs 2.5 Flash 💽 Enables [image]
Introducing Gemini 3.1 Flash-Lite, our fastest and most cost-efficient Gemini 3 series model. Built for high-volume workloads at scale, 3.1 Flash-Lite delivers high quality for its price and model tier. Rolling out in preview via Vertex AI → https://cloud.google.com/... [image]
Developers can now preview Gemini 3.1 Flash-Lite, our fastest and most cost-efficient Gemini 3 series model yet. With a 45% increase in output speed, it outperforms 2.5 Flash and features dynamic thinking levels to match task complexity. Rolling out in preview today in [video]
3.1 Flash-Lite outperforms 2.5 Flash with faster performance at a lower price. New ‘thinking levels’ let you dial in reasoning to adapt for different tasks, while still being able to handle complex workloads - like generating UI and dashboards or creating simulations. [image]
3.1 Flash-Lite is now live in the Gemini API 3.1 Flash-Lite outperforms 2.5 Flash across a majority of benchmarks, while being a little speedster! ⚡️⚡️⚡️ [video]
Gemini 3.1 Flash lite released and its next level price-performance-ratio Gemini 3.1 Flash-Lite, its fastest and most cost-efficient Gemini 3 model yet, built for high-volume developer workloads. Priced at just $0.25 per 1M input tokens and $1.50 per 1M output tokens, it [image]
Here are the benchmarks for the newly released Gemini 3.1 Flash-Lite Preview, now available in AI Studio. Good pricing @ $1.5/1M tokens on par with new Chinese OS models. It beats Qwen3.5 397B (~$3/1M) in shared benchmarks, so a great deal. It DOES NOT beat GLM-5 (~$2.5/1M). [ima…
socOCRbench has been updated to include the new Gemini 3.1 Flash-Lite, which achieves SOTA performance (by far) at this price point (and is best across all models for print tables!), as well as the new small Qwen 3.5 models, which are best in their respective size classes. [image…
Google has released Gemini 3.1 Flash-Lite Preview! This upgrades the fastest, lowest-cost Gemini model series, scoring 34 on the Artificial Analysis Intelligence Index while served at over 360 output tokens/sec, significantly faster than other first-party API endpoints Key [image…
Gemini 3.1 Flash-Lite was just released by @GoogleDeepMind, and the initial benchmarks are highly impressive for a model in this weight class. It is designed to be a high-volume, cost-sensitive workhorse that doesn't sacrifice quality. Based on the benchmark data, here is how it …
Excited to introduce Gemini 3.1 Flash-Lite, our fastest and most cost-efficient Gemini 3 series model designed for high-volume workloads. It's 2.5X faster Time to First Answer Token than 2.5 Flash (along with a 45% increase in output speed) ⚡️ [image]
Gemini 3.1 Flash-Lite is the fastest and most cost-efficient Gemini 3 series model⚡️ It outperforms 2.5 Flash with a 2.5X faster Time to First Answer Token and a 45% increase in output speed, at a fraction of the cost of larger models!
📢Introducing Gemini 3.1 Flash-Lite, our fastest and most efficient model, built for high-volume workloads. It outperforms 2.5 Flash in reasoning, reliability, and scalability at a lower cost. This model also introduces thinking levels. You can adjust compute by complexity of