DeepSeek moved V4-Flash output from $0.28 per million tokens to $1.32 at peak, with off-peak service at $0.66. The model does not get smarter by the hour; the queue gets tighter. By putting a clock on one token, DeepSeek exposed what customers are increasingly buying.

Key takeaways

  • DeepSeek’s V4-Flash output price rises from $0.66 per million tokens off-peak to $1.32 at peak, showing that congestion and delivery timing—not changing model intelligence—set the marginal price.
  • Open weights and cheaper architectures commoditize model access, but they do not provide reserved compute, predictable latency, burst capacity, or production reliability.
  • Agent schedulers can reduce costs by delaying flexible workloads, switching models, or paying premium rates only when immediate completion is necessary.
  • As model access gets cheaper, commercial value and infrastructure risk shift toward hosting, orchestration, power procurement, capacity planning, and reliable service.
  • DeepSeek’s reported losses and proposed infrastructure-funded capital raise illustrate how growing usage can create large financing requirements even when revenue rises quickly.

A weight file—or access to a capable model—captures less of the full cost of production use. A provider must supply throughput when customers want it, absorb bursts without failing, preserve latency under load, and keep enough capacity idle to honor an availability promise. The same model can be cheap even when immediate service is not.

That split changes what providers sell. They once priced tokens mostly as a tariff on capability: the smarter model cost more because intelligence was treated as the scarce input. DeepSeek’s rates instead charge for a constrained serving system at a particular moment.

Open weights commoditize the artifact, not the promise

Model makers are expanding the supply of usable intelligence from several directions. DeepSeek launched V4-Pro at $0.435 per million input tokens and $0.87 per million output tokens. Tencent released a 770-billion-parameter open model with a one-million-token context window.

Quarterly coverage volume: DeepSeekCoverage of DeepSeek by quarter, 2024 Q4 to 2026 Q3: from 4 to 62 articles per quarter, peaking at 133.peak 133622024 Q42026 Q3
Quarterly coverage · DeepSeek · 2024 Q4–2026 Q3 · current quarter projected

Each provider faces the same incentive. Open weights widen distribution, attract builders, and reduce the risk of exclusion from an ecosystem organized around proprietary agents. Falling architecture and access costs make release rational for multiple providers independently.

Publishing weights removes only one constraint. It does not provide a managed endpoint, reserve accelerators for a traffic spike, or guarantee that a million-token request finishes on schedule. Those jobs belong to the full serving system, where reliability, latency, utilization, and recurring production costs matter alongside benchmark quality.

Developers can move some inference away from centralized APIs through on-device models and self-hosted weights. That works for teams whose workloads fit the hardware and who can carry the operational burden. Managed serving remains valuable to buyers willing to pay someone else to keep the system running.

The clock now clears the token market

Nothing inside V4-Flash changes at the boundary between peak and off-peak hours. Customers simply pay more when their requests crowd the same serving capacity.

V4-Flash peak output price per 1M tokens
V4-Flash off-peak output price per 1M tokens

DeepSeek is applying compute-capacity economics through an API. At peak, one request occupies capacity that could serve another. Off-peak, the same request helps monetize infrastructure that would otherwise sit underused.

DeepSeek also charges flexible users half the peak rate, undercutting the idea that dynamic pricing is simply inflation with a nicer dashboard. Its tariff raises the cost of urgency and lowers the cost of delay, exposing congestion while giving movable demand a reason to move.

OpenAI and Anthropic have reported inference costs exceeding half of revenue. Training creates the model, but inference repeatedly produces what customers buy. Software margins become less software-like when each additional unit of revenue requires another pass through accelerators, memory, power, and networks.

OpenAI engineers reportedly found a method to more than halve inference cost. Broad deployment could ease pressure on token rates and improve margins at every utilization level.

But a cheaper serving system can still become congested when too many requests arrive together. Efficiency increases the work that fits into available capacity, while dynamic pricing shifts flexible work away from the queue.

Agents make delay an economic control

Time-of-use pricing works only when some demand can move. Interactive chat usually cannot because a person waiting for an answer treats latency as part of product quality. Agent workloads create a different demand shape: many tasks can continue without someone watching every token arrive.

For agentic inference, speed matters differently when humans are not directly in the loop. A background research task, code-validation run, document-processing batch, or multimodal workflow can often trade completion time for lower cost. The scheduler decides whether that trade is acceptable.

DeepSeek released an MIT-licensed agent harness with swappable model adapters and plugin-based components. Modular agent systems make model choice, tool choice, and execution policy programmable rather than fixed inside a chat interface.

Once an execution layer can see price and urgency, off-peak capacity becomes a product feature. An agent can route an immediate request to premium service, defer flexible work to a cheaper window, or move a task to another model. The useful metric becomes cost per completed task under a latency constraint, not the posted price per token.

DeepSeek’s experimental V4-Flash-Vision model bills images as tokens at V4-Flash pricing. The priced unit now covers visual input consumed inside an agentic workflow, making scheduling policy consequential across a larger share of machine work.

Agents operationalize the inference market’s capacity split. Cheap general capacity serves flexible work, while scarcer capacity earns a premium from workloads that require immediate response, predictable latency, or specialized performance. Software can make that procurement decision continuously.

Cheap access pulls model makers into infrastructure

Low model prices help providers win distribution, but they do not finance research or provision service on their own. Model companies must capture enough recurring inference spend to support both.

DeepSeek reportedly generated $70.7 million in revenue and a $106 million net loss during the first seven months of 2026, even as revenue reached roughly ten times its full-year 2025 level. Fast growth did not remove the financing requirement because service growth carries production costs with it.

DeepSeek responded by reportedly seeking $7.4 billion at a $74 billion valuation to fund research and expand into computing infrastructure. Cheap intelligence attracts usage; usage creates a capacity obligation; the obligation pulls the provider toward capital ownership.

Moonshot AI has pursued hosting agreements with Microsoft, Amazon, and Google for Kimi K3 while seeking up to a 30% revenue share. Rather than finance the entire serving layer itself, Moonshot is trying to capture more of the revenue generated by partners’ infrastructure.

Open-model providers can release more of the blueprint because hosting, orchestration, distribution, reliability, and customer access remain scarce. As the artifact becomes widely available, those operating capabilities carry more of the commercial value.

Token prices compress infrastructure risk

Providers that sell tokens assume exposure to compute, power, and network capacity. OpenAI is hiring a power-trading lead to hedge its data-center power portfolio, while the Shanghai Futures Exchange is exploring AI-token futures. Both moves treat inference demand as an operating risk that can be financed and hedged.

Dynamic API pricing is the retail edge of that structure. Customers see peak and off-peak rates. Providers manage hardware availability, electricity procurement, utilization, and the risk that contracted demand arrives before sufficient capacity does. A token price compresses that chain into one number.

Buyers then face three choices. They can delay flexible work to avoid congestion, pay a scarcity premium for immediate service, or move the workload to local or self-hosted deployment and assume the operating cost themselves.

A downstream buyer must now classify each workload before choosing a model: must it run now, can it wait, or can the team host it? Benchmark quality cannot answer those questions. The answers determine the bill and who carries the infrastructure risk.

DeepSeek’s $1.32 peak token and $0.66 off-peak token carry the same model intelligence. The hour changes the queue, the utilization, and who bears the infrastructure risk. AI competition is settling around scheduled, financed, reliably delivered capacity. The clock prices that promise.

DeepSeek’s August 2026 shift from model release to infrastructure financing

  • 2026-08-21 — DeepSeek unveiled an experimental V4 Flash model with visual-prompt understanding.
  • 2026-08-22 — DeepSeek unveiled an experimental multimodal version of V4 Flash.
  • 2026-08-26 — Sources reported $70.7 million in revenue and a $106 million net loss for the first seven months of 2026; full-year 2025 reportedly carried a $139 million net loss.
  • 2026-08-28 — DeepSeek was reportedly set to raise $7.4 billion at a $74 billion valuation and planned to expand into computing infrastructure.

Frequently asked questions

Why does the same V4-Flash output cost more at peak hours?

The model is unchanged, but peak requests compete for limited serving capacity. DeepSeek charges $1.32 per million output tokens at peak versus $0.66 off-peak, pricing urgency and congestion.

Does open-weight AI eliminate inference costs?

No. Open weights remove an access constraint, but production use still requires accelerators, memory, power, networking, workload scheduling, and operational reliability.

Which AI workloads can benefit most from off-peak pricing?

Deferrable agent tasks such as background research, code validation, document processing, and multimodal batches can trade completion time for lower cost. Interactive chat generally has less flexibility because users experience latency directly.

Should buyers compare models by token price alone?

No. The more useful measure is cost per completed task under a latency constraint, including whether work can wait, must run immediately, or can be self-hosted.

Why are low-cost model companies moving toward infrastructure?

Cheap access attracts usage, but serving that usage creates recurring capacity and reliability obligations. DeepSeek’s reported financial results and proposed funding for computing infrastructure show how model providers can be pulled toward capital-intensive operations.