/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

Sundar Pichai says Google is now processing 3.2 quadrillion tokens per month, up from 480T tokens per month a year ago and 9.7T tokens per month two years ago

and Doesn't Need You AnymoreCharles Rollet /Business Insider:Sundar Pichai announced at Google I/O that Gemini 3.5 Pro will launch next month; attendees groaned at the model coming out later than they expectedAlistair Barr /Business Insider:Google's latest AI numbers are huge. Here are the stats CEO Sundar Pichai just dropped.The Economic Times:Speed, cost, accessibility key in next phase of AI race, says Sundar PichaiThe Register:Google touts its tokenmaxxing and capex spending amid AI orgyFort

Axios Ina Fried

Context & Ripple Effects

Google’s earlier disclosures traced the same ramp: 1.3 quadrillion monthly tokens across its services in mid-2025, followed by reported growth in Gemini API activity and enterprise subscriptions. The new figure extends that trajectory from a product milestone to evidence of much heavier ongoing AI use across Google’s stack.

The coverage also shows that deployment is uneven: Gemini is gaining users and API traffic, while the broader Assistant-to-Gemini transition has slipped beyond Google’s prior timetable. Scale in model processing therefore does not automatically translate into completed product migration.

First-order effects

  • Google has a much larger operational AI workload to serve across Gemini, its direct API customers, and AI features embedded in its services; infrastructure capacity and inference efficiency become immediate execution priorities.
  • The reported increase strengthens Google’s case that Gemini has meaningful usage momentum as it prepares the next Pro-model release, even as some platform integrations remain delayed.

Second-order effects

  • Rivals competing for developers and enterprise AI workloads face a clearer need to match Google on serving scale, latency, and cost—not merely model launches—because API and consumer usage are growing alongside each other.
  • Google’s own product teams will face pressure to convert processing volume into reliable, broadly available features; the delayed Assistant migration highlights that integration, device rollout, and product readiness can remain constraints after compute is available.

Third-order effects

  • If this growth persists, AI competition will increasingly be organized around the economics of operating inference at mass-market scale, favoring firms that can pair models with large distribution channels and infrastructure.
  • Token counts will become a more common but imperfect operating metric: they indicate deployment intensity, while leaving open questions about revenue, user value, and the mix of internal versus external usage.

The trend: This is one data point in AI’s shift from benchmark-driven model competition toward a scale-and-efficiency contest for continuously serving models through consumer products and enterprise APIs.

Discussion

  • @google @google on x
    At last year's #GoogleIO, we were processing 480 trillion tokens a month across our surfaces. Now, we're processing over 3.2 *quadrillion* tokens a month. That's a 7x increase in just a year. These tokens represent problems being solved — by a user or a developer or a [image]
  • @edzitron Ed Zitron on x
    these guys will share literally any number other than “how much revenue we made on AI”
  • @mweinbach Max Weinbach on x
    Top companies in Google Cloud are doing 1T tokens a day, if they move 80% of their workloads to Flash they can save ~$1B annually according to @sundarpichai
  • @kimmonismus @kimmonismus on x
    This is inane. Tokens processed at insane scale! [image]
  • @officiallogank Logan Kilpatrick on x
    3.2 quadrillion tokens a month and still growing :) [image]
  • @alliekmiller Allie K. Miller on x
    Holy crap. And I don't say that lightly. Google is now processing 3.2 quadrillion tokens per month, up 7x from last year. That's a 3 with 15 zeroes after it. Actually, it's 3 and a 2 and 14 zeroes. #Google [image]
  • @jamespmcleod.ca James McLeod on bluesky
    “Token” is an annoying and stupid word, and I kinda hate how it's permeated society through both AI and crypto.  [embedded post]