Gemini 4 Argon supports up to 1M output tokens, up from 64K for prior models, and costs $2/1M input and $10/1M output tokens, but will rise to $4 and $20 later
Matthias Bastian /The Decoder:
Context & Ripple Effects
Google had been tuning Gemini’s Flash pricing through 2026: Gemini 3.6 Flash cut the price of 3.5 Flash, before Gemini 3.7 Flash positioned itself for coding and agent workloads at a lower launch rate. Argon breaks from that workhorse-price trajectory by attaching much larger output capacity to a materially higher output-token bill.
The release arrives alongside an Artificial Analysis comparison that places Argon High alongside GPT-6 Astra Max on its Intelligence Index, making the capacity and pricing terms consequential for buyers choosing high-end models rather than merely a SKU refresh.
First-order effects
- Developers using Gemini can produce up to 1 million tokens in one response rather than the prior 64K ceiling, enabling substantially larger single-run coding and agent outputs.
- Google’s introductory $2-per-million input and $10-per-million output pricing gives buyers a defined lower-cost entry window, while the stated move to $4 and $20 makes long-output workload budgeting more expensive afterward.
Second-order effects
- Teams that adopted Gemini 3.7 Flash for coding and agents must reassess whether Argon’s larger output allowance offsets its higher per-token cost, especially for output-heavy runs.
- GPT-6 Astra becomes a more direct procurement comparison for buyers seeking frontier-model capability, while Google differentiates Argon through output capacity and a staged price schedule rather than low-cost Flash positioning.
Third-order effects
- If high-end models keep expanding output limits while raising output-token prices, agent economics will hinge less on model access and more on controlling completion length, retries, and task decomposition.
- The split between lower-priced workhorse models and premium long-output models points toward a tiered inference market in which capacity limits and output pricing are product segmentation tools.
The trend: Frontier AI vendors are separating inexpensive workhorse inference from premium, long-horizon agent capacity, with output tokens becoming the central unit of both capability and cost.