Sundar Pichai says Google is now processing 3.2 quadrillion tokens per month, up from 480T tokens per month a year ago and 9.7T tokens per month two years ago
and Doesn't Need You AnymoreCharles Rollet /Business Insider:Sundar Pichai announced at Google I/O that Gemini 3.5 Pro will launch next month; attendees groaned at the model coming out later than they expectedAlistair Barr /Business Insider:Google's latest AI numbers are huge. Here are the stats CEO Sundar Pichai just dropped.The Economic Times:Speed, cost, accessibility key in next phase of AI race, says Sundar PichaiThe Register:Google touts its tokenmaxxing and capex spending amid AI orgyFort
Context & Ripple Effects
Google’s earlier disclosures traced the same ramp: 1.3 quadrillion monthly tokens across its services in mid-2025, followed by reported growth in Gemini API activity and enterprise subscriptions. The new figure extends that trajectory from a product milestone to evidence of much heavier ongoing AI use across Google’s stack.
The coverage also shows that deployment is uneven: Gemini is gaining users and API traffic, while the broader Assistant-to-Gemini transition has slipped beyond Google’s prior timetable. Scale in model processing therefore does not automatically translate into completed product migration.
First-order effects
- Google has a much larger operational AI workload to serve across Gemini, its direct API customers, and AI features embedded in its services; infrastructure capacity and inference efficiency become immediate execution priorities.
- The reported increase strengthens Google’s case that Gemini has meaningful usage momentum as it prepares the next Pro-model release, even as some platform integrations remain delayed.
Second-order effects
- Rivals competing for developers and enterprise AI workloads face a clearer need to match Google on serving scale, latency, and cost—not merely model launches—because API and consumer usage are growing alongside each other.
- Google’s own product teams will face pressure to convert processing volume into reliable, broadly available features; the delayed Assistant migration highlights that integration, device rollout, and product readiness can remain constraints after compute is available.
Third-order effects
- If this growth persists, AI competition will increasingly be organized around the economics of operating inference at mass-market scale, favoring firms that can pair models with large distribution channels and infrastructure.
- Token counts will become a more common but imperfect operating metric: they indicate deployment intensity, while leaving open questions about revenue, user value, and the mix of internal versus external usage.
The trend: This is one data point in AI’s shift from benchmark-driven model competition toward a scale-and-efficiency contest for continuously serving models through consumer products and enterprise APIs.