DeepSeek says it will lower V4 Pro API prices by 75% to $0.435/1M input and $0.87/1M output tokens, making permanent the discount prices set to expire on May 31
The prices listed below are in units of per 1M tokens.Vincent Chow /South China Morning Post:DeepSeek V4 Pro tops global bang-for-buck ranking after 75% price cutMarcus Schuler /Implicator.ai:DeepSeek Just Froze AI Prices at a Level Western Labs Cannot MatchAna Maria Constantin /The Next Web:DeepSeek made its 75% discount permanent. The AI price war just escalated.Jackson Chen /Engadget:DeepSeek permanently reduces the price of its flagship V4 model by 75 percentMinh Le /Tech in Asia:DeepSeek ma
Context & Ripple Effects
DeepSeek had already positioned V4 Pro and V4 Flash as low-cost offerings, while earlier off-peak discounts on V3 and R1 showed it was willing to use pricing to steer demand and compete. Making the V4 Pro reduction permanent extends that approach from a temporary promotion into its standard API economics.
The move comes as DeepSeek is reported to be expanding infrastructure, considering new financing and an eventual China IPO, and developing an inference chip. Those efforts make sustained low inference pricing strategically relevant rather than simply a short-term customer-acquisition tactic.
First-order effects
- Developers using V4 Pro receive a substantially lower ongoing cost base for input and output tokens, removing the May 31 expiry risk attached to the prior discount.
- DeepSeek accepts lower per-token revenue on its flagship API while strengthening the value proposition of V4 Pro relative to its own earlier pricing and other API choices.
Second-order effects
- Competing model API providers, particularly those selling on cost-performance, face greater pressure to justify higher token prices through capability, reliability, ecosystem integration, or matching discounts.
- Lower permanent pricing can shift more workloads from experimentation to production use, increasing demand for DeepSeek’s serving capacity and making infrastructure efficiency more consequential.
Third-order effects
- If providers can sustain sharp price cuts through efficient inference and greater infrastructure control, model APIs may compete increasingly on unit economics as well as model quality.
- DeepSeek’s reported work on an inference chip points to a possible longer-term split between providers that control more of their serving stack and those more exposed to third-party hardware costs; whether that translates into durable pricing advantage remains unproven.
The trend: This is part of an intensifying AI-inference price war in which providers seek to convert efficiency gains and infrastructure investment into permanently cheaper production APIs.