$200 a month still buys an unusually large effective token allowance from OpenAI and Anthropic. The price looks like abundance. The allocation says otherwise.

A fixed price now meets a variable bill

A premium subscription collects a fixed monthly fee. Inference creates a serving cost each time the product is used. More subscribers raise both revenue and serving costs. Heavier use by an existing subscriber raises only the latter.

That is the emerging inference economics problem. The commercial question is no longer only whether a frontier model can perform a task. It is whether the task merits frontier-model pricing every time it runs.

Customers are answering with routing. Companies are increasingly sending work to cheaper models, including offerings from Chinese providers. The cheapest model does not need to be the best model. It needs to be adequate for the work assigned to it.

Capability wins headlines. Cost per useful task increasingly decides the route.

OpenAI and Anthropic premium plans with very large effective usage allowances
Debt raised by Meta since 2022, roughly half of it in 2025

These numbers should not be divided. They belong to different companies and different accounting layers. The juxtaposition still matters.

One end of the market sells plentiful access at a stable retail price. The other finances the infrastructure beneath AI at enormous scale. Meta has also moved $30 billion of debt used to build AI data centers off its balance sheet through special-purpose vehicles. That does not isolate inference spending from training or other infrastructure costs. It does show how much capital the supply layer can absorb.

The customer sees $200. The system sees capacity allocation. The gap is the story.

Rationing and generosity serve the same strategy

Vendors are rationing supply while offering premium subscribers unusually generous token allowances.

These actions solve different sides of the same problem: preserving adoption without surrendering control of capacity. Large allowances build the habit. Rationing manages the bill created by that habit.

A subscription scale trap follows. A plan can become more attractive as its underlying service burden becomes harder to price. Success improves the adoption metric while worsening the allocation problem. Every local dashboard can remain green. The constraint simply moves downstream.

The question is not whether OpenAI and Anthropic want more usage. They remain willing to offer a great deal of it for $200. The question is which usage receives frontier capacity, which is redirected, and which eventually encounters a meter.

That is why customer substitution matters. When companies route work to cheaper models, they are not merely negotiating a lower price. They are decomposing demand. High-value tasks can retain expensive capacity. Routine tasks can move elsewhere. The model market begins to resemble an order book rather than a leaderboard.

The signals show pressure, not financial damage

The counter-case deserves precision. Generous $200 plans can preserve adoption even while usage costs rise. They may strengthen distribution, deepen customer habits, and keep competing models outside the workflow. High allowances are not evidence that the strategy has failed.

Nor do partner complaints about Anthropic prove that pricing pressure has reduced revenue or margins at either Anthropic or OpenAI. Complaints reveal bargaining tension. They are not audited unit economics.

The finding is narrower. Demand is becoming more price-sensitive. Vendors are rationing supply. Usage bills are rising rapidly. Premium access remains generous. These observations do not establish financial damage. They establish a changed constraint.

For much of the past year, the market treated model capability as the scarce asset. Today’s pricing signals place scarcity somewhere less glamorous: serving the model repeatedly at a price customers will continue to accept.

The best-model narrative can reverse with each release. The allocation machinery is harder to unwind. Once customers learn to route work by cost, every frontier request must justify its place against a cheaper substitute. Once vendors ration capacity, every generous allowance becomes an economic choice rather than a product flourish.

$200 is not proof that inference is cheap. It is a fixed numerator; routing and rationing determine how much frontier-model work sits beneath it. The price promises abundance. The allocation reveals its cost.