An evaluation by NIST's CAISI says DeepSeek V4 Pro lags behind leading US AI models by about eight months and is the most capable Chinese AI model to date
DeepSeek’s V4 launch followed a progression from V3 and V3.1, where the company emphasized benchmark gains and adaptation to next-generation Chinese-made chips. Earlier reporting also framed export controls as a constraint that pushed the lab toward more resource-efficient development.
The company previewed V4 Pro and V4 Flash while positioning V4 Pro as several months behind the frontier. CAISI’s evaluation independently narrows that claim to roughly an eight-month gap while identifying it as China’s strongest model so far, giving buyers a clearer relative-performance reference.
First-order effects
DeepSeek gains a credible external marker as the leading Chinese model provider, but the evaluation also establishes that V4 Pro remains behind leading US systems on CAISI’s assessment.
Enterprise and public-sector model evaluators can treat DeepSeek’s own positioning with more discipline: Chinese leadership and global-frontier parity are distinct conclusions.
Second-order effects
US model providers retain a performance-based sales argument against DeepSeek, while Chinese competitors face a higher bar to displace it as the domestic capability leader.
The result increases the value of workload-specific procurement: a model can be the strongest locally available option without being the best choice for tasks that require frontier-level performance.
Third-order effects
If independent evaluations become a regular reference point, model competition will be judged less by launch claims and more by repeatable comparisons across capability, deployment constraints, and fit for particular workloads.
The gap between national model ecosystems may increasingly matter alongside absolute benchmark leadership, as buyers balance access, sovereignty, and performance rather than treating AI models as interchangeable.
The trend: AI procurement is shifting from headline benchmark competition toward independently assessed, sovereignty-aware selection among increasingly capable regional model ecosystems.
DeepSeek V4 has a similar capability to GPT-5, released 8 months ago, according to a new @NIST report. If the current trend continues, we'll see a Chinese model at GPT-5.5 (roughly Mythos-level) model around February 2027. [image]
In April 2026, the Center for AI Standards and Innovation (CAISI) evaluated the open-weight AI model DeepSeek V4 Pro ("DeepSeek V4"). CAISI evaluations indicate that DeepSeek V4's capabilities lag behind the frontier by about 8 months https://www.nist.gov/...
New composite eval of DeepSeek V4 from CAISI suggests China is falling behind. Notice the relative steepness of their improvement trend. https://www.nist.gov/... [image]
The issue is that benchmarks simultaneously undersell and oversell the gap. DeepSeekv4 belongs to an entirely new category of model by side and by design, and the most dramatic step taken this year to bring the open architecture ecosystem closer to frontier.
This is a good explanation of why the gap between open and closed models is larger than it appears in benchmarks. I would add in that current open models are also more fragile than closed: they handle out-of-distribution problems far less well & have lower emergent capabilities.