Artificial Analysis estimates DeepSeek V4-Flash's average cost at $0.03 per test, compared with Kimi K3's $0.86, GPT-5.6 Sol's $1.86, and Claude Fable 5's $3.15
A version of Chinese startup DeepSeek's flagship AI model is by far the least expensive to run on benchmark tests among …
Context & Ripple Effects
DeepSeek had already positioned V4 Flash at $0.14 per million input tokens and $0.28 per million output tokens in April; the new comparison translates those API rates into a concrete benchmark-run cost. Its reported Intelligence Index score matching Gemini 3.6 Flash gives the cost comparison a quality reference point.
The result also extends a pricing pattern: DeepSeek had made a 75% V4 Pro API price cut permanent in May. The relevant competitive question is increasingly cost at a given level of measured capability, rather than token pricing in isolation.
First-order effects
- DeepSeek V4-Flash becomes the lowest-cost option in this benchmark comparison at $0.03 per test, well below the cited costs for Kimi K3 and GPT-5.6 Sol.
- Teams using these benchmark tests can evaluate V4-Flash at materially lower run cost, while Kimi and GPT-5.6 Sol face an immediate price-performance disadvantage in this comparison.
Second-order effects
- Model providers competing for cost-sensitive API workloads may face pressure to lower inference prices or demonstrate capabilities that justify higher benchmark-run costs.
- Buyers are likely to put more weight on cost-per-test and similar workload-level measures, reinforcing the value of the previously disclosed V4 Flash token pricing as an input to procurement decisions.
Third-order effects
- If comparable quality can be sustained at sharply lower inference cost, model competition may shift further from headline model capability toward operational efficiency and price-performance.
- Benchmark economics will become more consequential, but results will remain sensitive to which tests are used and whether benchmark performance maps to a buyer's production workload.
The trend: This is another data point in the shift toward competing on effective inference cost per useful AI task, not simply model capability or nominal token prices.