The falling price of a token and what it unlocks.
Inference cost is the recurring expense of running an AI model for users, making it central to API pricing, product margins and the range of AI tasks that can be deployed at scale. Per-token prices have fallen, but expanding reasoning workloads can raise total spending by consuming more tokens, compute and time. The result is a market in which model capability, serving efficiency, infrastructure and pricing strategy are increasingly intertwined.
Training builds a model, but inference is the continuing cost of serving prompts, generated outputs and production workloads. Providers must recover these costs through usage pricing, subscriptions, enterprise contracts or other monetization, while also supporting reliable delivery and integration. This makes AI unit economics structurally different from software businesses whose marginal cost of serving another user may be much lower.
For buyers, the relevant measure is increasingly cost per useful task rather than a headline model price alone. That includes the tokens used, the reliability of the result, required human oversight and the value produced in a workflow. A low-cost model can be attractive for routine work, while a more expensive model may be justified where stronger performance reduces downstream effort or errors.
The price per token for AI models has declined, and providers have continued to introduce materially different price tiers across model families. OpenAI, for example, has offered models with distinct input and output token prices, while DeepSeek has publicized low-priced V4 variants. Such pricing makes model selection and workload routing a more active part of product design and procurement.
Lower unit prices do not automatically mean lower bills. The Wall Street Journal reported that newer reasoning models can require more tokens to complete tasks, raising developer costs despite falling per-token prices. Test-time compute extends this tension: spending more compute during a task can improve performance, but turns reasoning into a variable operating expense rather than a fixed feature.
Serving efficiency can alter both provider margins and the prices available to customers. Reports that OpenAI engineers found a way to more than halve inference costs illustrate why improvements in the serving stack can be commercially consequential even without a change in the customer-facing task. Smaller models are also characterized as faster and cheaper options for many business tasks, creating incentives to match capability to workload rather than use the largest model everywhere.
Efficiency is not limited to the model itself. Specialized inference offerings from companies such as Cerebras and Groq, cloud capacity, chips, data centers and distribution channels can all affect the marginal cost and latency of serving a model. This turns inference into strategic infrastructure: providers that can supply capacity efficiently and reliably may have more room to lower prices or protect margins.
Inference can absorb a large share of AI revenue. The Financial Times reported that OpenAI and Anthropic presented profitability projections with and without training costs and reported inference costs exceeding half of revenue, while Anthropic projected a lower gross margin after higher inference costs. These dynamics show why revenue growth alone does not settle the economics of a model provider.
Competition can compound the squeeze. Companies facing rising AI costs have increasingly used cheaper models, including models from China, putting pricing pressure on OpenAI and Anthropic. Lower-cost Chinese models from companies including DeepSeek and MiniMax have also gained token consumption relative to US rivals, according to OpenRouter data cited by the Financial Times, strengthening buyers' alternatives and switching leverage.
Cheaper inference expands the set of workloads that can clear an economic threshold, particularly high-volume, latency-sensitive or continuously used applications. It can support broader experimentation and shift competitive advantage away from model quality alone toward workflow integration, enterprise distribution, pricing and operational reliability. Benedict Evans has described a direction in which frontier models become commodity infrastructure as token constraints ease and value shifts toward products built on top.
The central management question is not simply how to minimize compute, but how to budget it according to task value and error tolerance. Adaptive inference budgeting can allocate tokens, latency, tool calls and model capacity where additional computation is likely to matter, while cheaper or smaller models handle simpler work. Key indicators are the relationship between per-token prices and tokens consumed, provider inference costs as a share of revenue, the availability of lower-cost substitutes, and whether product-level value rises faster than recurring serving costs.
Grounded in the archive and knowledge graph. Browse all topic guides, the concept reference, or the posts.