An analysis of 100T+ tokens from the past year shows reasoning models now represent over half of all usage, open-weight model use has grown steadily, and more
this is not a model I hear much about. [image] @openrouterai : We collaborated with @a16z to publish the **State of AI** - an empirical report on how LLMs have been used on OpenRouter. After analyzing more than 100 trillion tokens across hundreds of models and 3+ million users (excluding 3rd party) from the last year, we have a lot of [image] @a16z : >100 trillion token analysis of reasoning model usage over time Full piece from @MaikaThoughts, @AnjneyMidha, @xanderatallah, and @cclark: https://openrouter.ai/... [image] @scaling01 : The moment open-source models were close to 30% of OpenRouter traffic and almost all of them came from China with the notable models being: DeepSeek V3/R1, Qwen3 family, Kimi-K2 and GLM-4.5 + Air Minimax M2 is now also a major player, but open-weights models token-usage [image] Nathan Lambert / @natolambert : On a prompt count basis this mean reasoning models are not close to a majority on OpenRouter, as reasoning models can use 10-1000x the tokens of non-thinking models per prompt. Lots of need for fast, efficient open models. Reasoning model usage is likely closed labs more. [image] Bluesky: Tim Duffy / @timfduffy.com : Lots of interesting details in this new report on usage trends from OpenRouter. openrouter.ai/state-of-ai I've been wondering about mean coding input token length, in their data it's around 20k tokens. Other large categories (roleplay, technology science) average around 5k [image]
OpenRouter
Context & Ripple Effects
OpenRouter’s dataset turns the shift toward reasoning into an observable usage pattern: these models take a majority of tokens, but not prompts, because they can use far more tokens per request. That distinction matters for evaluating demand and cost, rather than treating token share as a simple user-preference measure.
Model providers and OpenRouter users must treat reasoning inference as the main source of token consumption, with efficiency and latency becoming immediate product considerations.
Open-weight models gain stronger evidence of practical demand on a multi-model platform, while their growing share increases the set of viable alternatives for users.
Second-order effects
Routing layers and application developers have greater incentive to select models by task and token economics, especially for long-context coding workloads that average roughly 20,000 input tokens in the dataset.
Closed-model vendors face more pressure to compete not only on reasoning quality but on the cost and deployability of alternatives; the reported traffic was concentrated in Chinese open-source models.
Third-order effects
If token-intensive reasoning remains the dominant workload, AI competition is likely to shift toward cost per useful task and inference capacity, not just benchmark leadership.
Steady open-weight adoption could make model access more geographically and commercially diverse, although OpenRouter traffic alone cannot establish the broader market mix.
The trend: AI usage is moving toward a multi-model inference market in which reasoning capability, open-weight availability, and serving economics are increasingly inseparable.
this one chart explains EVERYTHING about why OpenAI, xAI and Deepmind dropped everything to go chase after the grand prize in koding usecases as i said at AIE CODE and in my cogpost, Code AGI will be achieved in 20% of the time of full AGI, and capture 80% of the value of AGI. [i…
quite rich report from openrouter. points worth caring: - oss models have grown to have roughly 30% share on openrouter - code and companionship still major use-cases, ~70-80% - remaining use cases are like translation, trivia, general knowledge questions - chinese models have [i…
@AnjneyMidha ... One finding: we observe a Cinderella “Glass Slipper” effect for new models. Early users a new LLM either churn quickly or become part of a foundational cohort, with much higher retention than others. They are early adopters who can “lead” the rest of the market (…
OSS isn't “just for tinkering” - it is extremely popular in two areas: 🧙♂️ Roleplay / creative dialogue: >50% of OSS usage 🧑💻 Programming assistance: ~15-20% [image]
We hope that it will help you make better decisions about where to invest, what to build, and where AI adoption is heading next. Read the paper here https://openrouter.ai/...
Most interesting chart in this for me so far is this one: I'm surprised about the the large MiniMax M2 share—this is not a model I hear much about. [image]
We collaborated with @a16z to publish the **State of AI** - an empirical report on how LLMs have been used on OpenRouter. After analyzing more than 100 trillion tokens across hundreds of models and 3+ million users (excluding 3rd party) from the last year, we have a lot of [image…
>100 trillion token analysis of reasoning model usage over time Full piece from @MaikaThoughts, @AnjneyMidha, @xanderatallah, and @cclark: https://openrouter.ai/... [image]
The moment open-source models were close to 30% of OpenRouter traffic and almost all of them came from China with the notable models being: DeepSeek V3/R1, Qwen3 family, Kimi-K2 and GLM-4.5 + Air Minimax M2 is now also a major player, but open-weights models token-usage [image]
On a prompt count basis this mean reasoning models are not close to a majority on OpenRouter, as reasoning models can use 10-1000x the tokens of non-thinking models per prompt. Lots of need for fast, efficient open models. Reasoning model usage is likely closed labs more. [image]
Lots of interesting details in this new report on usage trends from OpenRouter. openrouter.ai/state-of-ai I've been wondering about mean coding input token length, in their data it's around 20k tokens. Other large categories (roleplay, technology science) average around 5k [imag…