Chinese models now account for nearly 60% of token usage by US companies on OpenRouter. A US application can create that share without handing its entire workload to any one Chinese vendor, because software now chooses the supplier one request at a time.

Key takeaways

  • Inference is shifting from vendor commitment to request-level procurement: applications can choose among models for each call based on quality, latency, availability, and price.
  • OpenRouter’s scale—25 trillion tokens weekly across more than 400 models—shows that routing can create a functioning market rather than merely simplify API access.
  • Chinese open-weight models provide credible low-cost substitutes, accounting for nearly 60% of US-company token usage on OpenRouter and increasing buyers’ leverage over proprietary suppliers.
  • Cost reductions no longer automatically become provider margin when routers can redirect traffic; buyers capture more savings unless a model remains irreplaceable for a valuable workload.
  • The durable control point is the evaluation and workflow layer, but routers must outperform buyers’ internal systems through reliable selection, fallbacks, compliance, and accountability.

Buyers capture a falling cost only when they can switch suppliers. A long model menu does little on its own. Applications gain leverage when software can choose repeatedly according to the quality, latency, availability, and price required for each call.

An application attached to one model leaves the provider to decide how much of an efficiency gain to keep. When the application can route among credible substitutes, each provider has to compete for the workload again. A shared interface creates that competition without supplier coordination.

A router converts an API call into a sourcing event

OpenRouter’s weekly volume grew fivefold in six months, from 5 trillion to 25 trillion tokens, across more than 400 models.

tokens routed weekly across 400+ models

At that scale, routing goes beyond developer convenience. Software can assign each request to an expensive frontier model, a cheaper substitute, a fallback supplier, or several models in parallel. Applications can make model choice an operating decision thousands of times a minute instead of an architectural commitment revisited every six months.

Meta’s behavior is more revealing than any routing company’s pitch. The company is developing Switchboard to direct some tasks toward lower-cost models rather than paying top-model prices indiscriminately. Meta can afford frontier inference, but its scale makes segmentation valuable: the relevant question becomes which model is sufficient for this request.

Developers first used cheaper inference to run the same applications for less. Routers let them redesign applications around different grades of inference for different jobs, making the request the unit of competition.

Chinese open-weight supply made substitution operational

Routers need credible substitutes. Chinese open-weight models have supplied them at enough quality and low enough cost to move diversification from a vendor-risk plan into production.

Chinese models account for nearly 60% of token usage by US companies on OpenRouter. Lower-cost models from DeepSeek and MiniMax began overtaking US rivals in token consumption in February. Those consumption figures record assigned workloads rather than benchmark votes.

Inference clouds make that substitution easier. Cursor accesses Moonshot AI’s Kimi K2.5 through Fireworks AI for its Composer 2 integration. The US application does not need to reorganize itself around a Chinese supplier because Fireworks absorbs the integration boundary and supplies the model as another source of capacity.

Open weights enlarge the supplier pool even when buyers do not host the models themselves. Routers, inference clouds, and managed platforms can offer more substitutes without forcing each application to rebuild its product.

Regulators and corporate security teams can narrow that pool. Restrictions on Chinese models would disrupt workloads that already use them, while compliance rules can disqualify a technically substitutable model. Even within those limits, US applications selected lower-cost Chinese models once inference clouds made them easy to insert.

Quality is becoming a portfolio result

One model sets an application’s quality ceiling only when it receives every task. Routing breaks that bundle. Routine requests can go to inexpensive models, difficult requests to frontier systems, and failed requests to a second supplier. Applications can also combine several outputs for complex work.

OpenRouter’s Fusion makes the portfolio logic explicit by prompting multiple models in parallel and synthesizing the results. The company claims Fusion can reach or surpass Fable-level performance on deep-research tasks at half the price.

The claim does not settle the economics. Fusion reportedly kept a premium Opus model in the call stack as a judge. Buyers must count the full stack because an expensive evaluator can absorb the savings from cheaper underlying models.

Applications can optimize an inference portfolio across cost, latency, and performance. Buyers no longer ask only which model ranks highest on a standalone benchmark; they ask how cheaply the workflow can achieve an acceptable result.

The cheapest model can set the price for routine work even when the best model sets the quality ceiling.

Frontier models can retain substantial pricing power for the hard tail. Routing may even concentrate premium demand by reserving the best model for calls where its advantage is measurable and valuable. The broader inference market is splitting between cheap general capacity and deliberately selected specialist capacity, with premium economics reserved for models whose removal damages the outcome enough to justify their price.

Supplier efficiency no longer guarantees supplier margin

Model providers still have powerful ways to lower their costs. OpenAI and Anthropic have told investors that inference costs exceed half of revenue, making serving efficiency central to their economics. OpenAI engineers also reportedly found a method to more than halve inference cost.

A provider can use such gains to lower prices, raise margins, or support more usage. When several providers can serve a request and applications can move among them, however, one supplier’s cost reduction becomes a new bid in a competitive market. Rivals respond, routers reallocate traffic, and part of the gain passes to the buyer.

Suppliers may expect cheaper tokens to stimulate enough demand for them to capture an expanding market. Total demand can still rise, but no provider is guaranteed the resulting value. Buyers can segment demand, reserve premium models for high-value calls, and force routine work into a lower-priced pool.

For buyers, the operating response is concrete: measure outcomes by task, set cost and latency ceilings, and reserve frontier models for requests where cheaper systems fail. Without trustworthy evaluations, routing amounts to indiscriminate cost-cutting.

The control plane has to earn its position

Investors are pricing access, customization, and selection as substantial commercial functions. OpenRouter passed $50 million in annualized revenue, while Fireworks AI reportedly sought funding at a $15 billion valuation after being valued at $4 billion in October 2025.

Routers cannot assume that position is defensible. Meta’s Switchboard shows that a large buyer can internalize routing once its inference bill and workload volume justify the effort. Large buyers can replicate a common endpoint, so routers need a deeper advantage than access.

A router retains value only when it makes better decisions than the buyer can make alone. It must provide trustworthy evaluations, reliable fallbacks, latency management, compliance, and integration with the workflow that defines success. Its moat lies in choosing the model for the next consequential call, explaining the choice, and bearing responsibility when it fails.

At OpenRouter’s scale, that nearly 60% share is buyer leverage made visible. Across 25 trillion weekly tokens, credible substitutes, freedom to switch, and trusted evaluations can turn each request into a fresh bid, with the application deciding who wins.

Evidence that inference is becoming a routed market

Evidence dateMarket signalSpecific figure
2026-05-27OpenRouter weekly volume25 trillion tokens, up from 5 trillion six months earlier
2026-05-27Available supplier pool400+ models
2026-06-15Claimed economics of OpenRouter FusionFable-level intelligence at half the price
2026-07-22Chinese-model share of US-company usage on OpenRouterApproximately 60% of tokens

Frequently asked questions

How does model routing turn inference into a procurement market?

A router evaluates each request and assigns it to a suitable supplier rather than binding the application to one model. That forces providers to compete repeatedly for workloads on price, quality, latency, and availability.

Why are Chinese open-weight models important to this shift?

They expand the pool of credible, lower-cost substitutes: Chinese models account for nearly 60% of US-company token usage on OpenRouter. Inference clouds and managed platforms let applications use those models without integrating directly with each vendor.

Will frontier models lose all pricing power?

No. Routing can reserve frontier models for difficult or high-value requests where their removal measurably harms results, while cheaper models set the price for routine work.

Can multi-model workflows always lower costs?

No. OpenRouter claims Fusion can deliver Fable-level deep-research performance at half the price, but its reported use of a premium Opus model as a judge shows why buyers must evaluate the cost of the entire workflow.

What makes a routing platform defensible?

A common endpoint alone is replicable, especially by large buyers such as Meta. A router needs superior task evaluations, dependable fallbacks, latency and compliance controls, workflow integration, and accountability for consequential routing decisions.