DeepSeek V4 Pro costs $1.74/1M input and $3.48/1M output tokens while V4 Flash costs $0.14/1M input and $0.28/1M output tokens, both the cheapest in their class
Chinese AI lab DeepSeek's last model release was V3.2 (and V3.2 Speciale) last December. They just dropped the first of their …
Context & Ripple Effects
DeepSeek’s V4 launch pairs a much larger Pro model and a smaller Flash model with a 1M-token context window, extending the lab’s V3.2-era lineup into distinct performance and cost tiers. The stated pricing positions both tiers at the low end of their respective classes.
The move follows DeepSeek’s earlier claim of unusually favorable V3 and R1 inference economics. Subsequent coverage says the company made a further V4 Pro discount permanent, making this launch price a starting point in an ongoing pricing strategy rather than a one-off benchmark.
First-order effects
- Developers can immediately choose between a lower-cost V4 Flash API for price-sensitive workloads and V4 Pro for more capable workloads, with token charges explicitly separated for input and output.
- DeepSeek sets a low public price anchor for long-context model usage while trying to monetize two different model-size tiers.
Second-order effects
- Competing API vendors face sharper procurement comparisons on both token pricing and usable capability, especially for workloads with large prompts or outputs.
- Enterprise buyers gain more reason to route workloads by model tier and measured task performance rather than standardize on a single premium API; DeepSeek’s later Pro cut reinforces that price may remain negotiable or volatile.
Third-order effects
- If low-priced frontier-adjacent models remain available, inference economics—not only headline model quality—will increasingly determine which providers win production workloads.
- The pattern could shift model vendors toward segmented portfolios and recurring price reductions, while placing greater importance on the infrastructure efficiency behind those prices; whether DeepSeek can sustain that structure depends on its operating economics and capacity buildout.
The trend: This is part of the shift from a scarcity-priced AI API market toward performance-tiered inference competition in which cost per useful workload becomes a primary buying criterion.