DeepSeek details V3.1 and says it surpasses R1 on key benchmarks and is customized to work with next-gen Chinese-made AI chips, after unveiling it on August 19
DeepSeek unveiled an update to an older model that it says surpasses the seminal R1 on key benchmarks, keeping the Chinese startup …
Bloomberg
Context & Ripple Effects
V3.1 arrived first as a sparse release with a longer context window; this update supplies the missing performance and hardware-positioning details around the August 19 V3.1 launch.
The claim extends DeepSeek’s recent effort to improve R1’s math, programming and logic performance, while shifting attention from model quality alone to where the model can run: its earlier R1 update set that performance arc.
First-order effects
DeepSeek can present V3.1 as both an upgrade over R1 on selected benchmarks and a model tailored for next-generation Chinese-made AI chips, giving users a more specific deployment option.
Chinese AI-chip makers gain a named model-compatibility claim that can help validate their platforms for customers evaluating inference deployments.
Second-order effects
Model providers and chip vendors serving the same market face pressure to show comparable software optimization and benchmark performance, not merely raw hardware specifications.
Buyers may place greater weight on tested model-chip combinations when choosing AI infrastructure, favoring suppliers able to support an integrated stack.
Third-order effects
If such co-optimization becomes routine, AI competition can increasingly turn on full-stack deployment—models, inference software and locally available chips—rather than frontier-model claims in isolation.
That shift could broaden the role of second-source compute in procurement, though DeepSeek’s benchmark and compatibility assertions alone do not establish broad customer adoption.
The trend: This is one data point in the move toward regionally integrated AI stacks, where model advances are paired with optimization for alternative compute supply.
new deepseek. if it's a top tier coding model as rumored i am probably abandoning western labs for the forseeable future huggingface.co/deepseek-ai/...
Introducing DeepSeek-V3.1: our first step toward the agent era! 🚀 🧠 Hybrid inference: Think & Non-Think — one model, two modes ⚡️ Faster thinking: DeepSeek-V3.1-Think reaches answers in less time vs. DeepSeek-R1-0528 🛠️ Stronger agent skills: Post-training boosts tool use and
DeepSeek launches V3.1, unifying V3 and R1 into a hybrid reasoning model with an incremental increase in intelligence Incremental intelligence increase: Initial benchmarking results for DeepSeek V3.1 show Artificial Analysis Intelligence Index of 60 in reasoning mode, up from [im…
DeepSeek-V3.1 benchmarks just dropped and... holy efficiency batman 🦇 • 66% on SWE-bench (best open model) • 5.5x faster at terminal tasks than R1 • 3.4x better at web browsing tasks • Still just 37B active params per token This is what happens when you solve hybrid AI [image]
DeepSeek v3.1 just dropped. Key highlights: - Big improvements on coding, agentic and reasoning vs older Deepseek R1. Open source SOTA - Still slightly behind other closed SOTA models in benchmarks. E.g. 66.0% on SWE-Bench Verified vs GPT-5's 74.9% and Opus 4.1's 74.5% - [image]
API Update ⚙️ 🔹 deepseek-chat → non-thinking mode 🔹 deepseek-reasoner → thinking mode 🧵 128K context for both 🔌 Anthropic API format supported: https://api-docs.deepseek.com/ ... ✅ Strict Function Calling supported in Beta API: https://api-docs.deepseek.com/ ... 🚀 More API resour…
And so we know what V3.1 is. Yes, it's an agent. - continued long-context pretrain for 840B tokens - *significant* gains in agentic regimes (I've hopefully accurately aggregated some tables) They responded to GLM&Kimi. ...They didn't announce V4. [image]
Model Update 🤖 🔹 V3.1 Base: 840B tokens continued pretraining for long context extension on top of V3 🔹 Tokenizer & chat template updated — new tokenizer config: https://huggingface.co/... 🔗 V3.1 Base Open-source weights: https://huggingface.co/... 🔗 V3.1 Open-source weights: