DeepSeek V4 Pro has 1.6T parameters, DeepSeek's largest model by that metric, and V4 Flash has 284B parameters; both models have a context window of 1M tokens
South China Morning Post:
Context & Ripple Effects
DeepSeek’s V3 established a large open-source MoE baseline at 671B total parameters, while the V4 launch separates its offering into a high-capacity Pro model and a smaller Flash model.
Related coverage places the launch alongside unusually low listed token prices and an acknowledged gap between V4 Pro and the state of the art. That makes the release as much a serving-and-product segmentation move as a raw model-scale milestone.
First-order effects
- DeepSeek now offers customers a choice between a much larger flagship and a lower-capacity variant while giving both access to a 1M-token context window.
- The larger Pro model raises DeepSeek’s disclosed model-scale ceiling; Flash provides a lower-priced route into the same long-context product family.
Second-order effects
- Low pricing for both tiers puts pressure on rival model providers to justify premiums through measurable capability, reliability, or serving performance rather than parameter counts alone.
- A 1M-token window increases the importance of inference efficiency and available serving capacity, since long-context access can become costly to operate even when token list prices are low.
Third-order effects
- If this two-tier pattern persists, frontier-model competition will increasingly be organized around workload segmentation: expensive capacity for demanding tasks and aggressively priced models for broad deployment.
- The durable constraint shifts toward infrastructure and inference economics: long-context models can expand buyer choice, but only providers that can serve them efficiently can sustain that choice at low prices.
The trend: This is one data point in the shift from single flagship-model launches toward tiered, long-context AI offerings differentiated by inference cost and operational capacity.