DeepSeek debuts DeepSeek-V4.1-Flash, its smallest model built on its new Causal Encoder-Decoder architecture, with 552B backbone parameters and 1M-token context
Chinese artificial intelligence startup DeepSeek on Thursday launched DeepSeek-V4.1-Flash, which the company said is the smallest model …
Reuters
Context & Ripple Effects
DeepSeek had already separated its V4 line between a 1.6T-parameter V4-Pro and a 284B-parameter Flash model, both with 1M-token context. Its earlier V4 lineup paired very large backbones with long-context processing, while V4-Pro was positioned as a lower-priced competitor on selected benchmarks.
The company also tested multimodal V4 Flash capabilities in August. V4.1-Flash makes the architecture itself the next point of differentiation, as DeepSeek has said demand and AI-agent workloads are straining its backend systems.
First-order effects
- DeepSeek adds a 552B-parameter V4.1-Flash as the smallest member of its Causal Encoder-Decoder family, extending the V4 portfolio beyond the prior Flash design while retaining a 1M-token context window.
- Developers with long-document or agentic workloads gain a new DeepSeek model option that does not require using the company’s 1.6T-parameter V4-Pro.
Second-order effects
- DeepSeek’s lower-priced V4-Pro positioning already put pressure on Kimi K3 in benchmark-led buying decisions; a new Flash-tier architecture broadens that pressure to customers weighing long-context capability against model size.
- Serving efficiency becomes more consequential for DeepSeek because it is simultaneously hiring senior engineers to address backend strain tied to user demand and AI-agent complexity.
Third-order effects
- The V4 line suggests frontier model competition is becoming less about maximizing total parameters alone and more about architectures that make long-context, multimodal and agent-oriented workloads practical to serve.
- If model buyers reward those trade-offs, providers with differentiated inference architectures may compete more directly with larger general-purpose models without matching their backbone size.
The trend: Frontier AI vendors are segmenting model portfolios around architecture and serving efficiency, not simply building ever-larger parameter counts.
Related: Frontier minimum efficient scale · Embedded AI agents · DeepSeek · Causal Encoder-Decoder · DeepSeek’s experimental multimodal V4 Flash · DeepSeek’s V4 model lineup
Related Coverage
- DeepSeek-V4.1-Flash — Image-Text-to-Text Transformers Safetensors deepseek_v41 text-generation 8-bit precision fp8 Hugging Face
- DeepSeek v4.1 Flash Hacker News
- China's DeepSeek Launches Smaller, Faster AI Model Caixin Global · Yang Zirui
- DeepSeek's New Low-Cost Model Deals a Fresh Blow to OpenAI, Z.ai Bloomberg · Saritha Rai
- New Deepseek model V4.1-Flash cuts memory needs for AI agents The Decoder · Jonathan Kemper
- DeepSeek V4.1 Flash Beats OpenAI's GPT-5.6 Sol And Anthropic's Opus 5 On Coding And Cybersecurity At An ~86x Lower Cost, While Reducing HBM Requirements By 3.8x And SSD Ones By 8x Wccftech · Rohail Saleem
- DeepSeek Unveils V4.1-Flash Model with Architectural Upgrades, Price Cuts Ahead of Shanghai IPO Techstrong.ai · Jon Swartz
- DeepSeek's New Model Nearly Matches GPT-6 Astra on Design—at 1.4% of the Cost Decrypt · Jose Antonio Lanz
Discussion
-
@deepseek_ai
@deepseek_ai
on x
🚀 Introducing DeepSeek-V4.1-Flash: smarter, faster, more efficient. 🔹 Introducing the smallest model in our new architecture family, with native visual understanding. 🔹 Designed for greater capability, faster inference, higher throughput, and scaling to larger models. 1/6
-
@kimmonismus
@kimmonismus
on x
DeepSeek just released V4.1-Flash with a new architecture, six weeks after its July V4-Flash update. July's release improved post-training while keeping the architecture unchanged. (Same with GLM-5.3/Flash) V4.1 introduces a Causal Encoder-Decoder architecture with native visual …
-
@victor207755822
Deli Chen
on x
Just Try It !
-
@sudoingx
@sudoingx
on x
deepseek keeps punching above its weight man, this is insane. the new v4.1 flash is a 552B MoE with a new asymmetric trick, 8B active parameters on input, 16B on output, and look what that buys in the table. it beats their own v4 PRO on almost every agentic bench, terminal bench …
-
@rishdotblog
Rishabh Srivastava
on x
Wow. Cheaper than luna (atleast on a per token basis), with perf between terra and sol Thank you deepseek - the openness with which they're sharing their IP is incredible! Amazing to see the price-performance frontier improve so rapidly
-
@doubilitysteven
Wen Liu
on x
make “encoder-decoder” great again😉
-
@aravsrinivas
Aravind Srinivas
on x
Wow!
-
@jenzhuscott
Jen Zhu
on x
Instead of getting anxious by brute compute agent swarm solving a math problem you've never heard of, or an unknown researcher's doom post... Friends, here is a small/mighty OSS model that gives you can customize while costs you nothing. No dooms day drama, no lecturing, no virtu…
-
@theahmadosman
Ahmad
on x
You see this? This is why I told you all to go get yourselves some compute Models will continue to get better, more intelligent, capable, efficient, and smaller
-
@deepseek_ai
@deepseek_ai
on x
💾 Smaller KV cache. Bigger savings. Compared with the previous generation, V4.1-Flash's KV cache needs just: 🔹 1/4 the HBM 🔹 1/8 the SSD storage Cache-hit charges often account for a large share of agent costs. Compressing the cache cuts those costs significantly. 3/6
-
@zizhpan
Zizheng Pan
on x
DeepSeek-V4.1-Flash is now publicly available on App, Web, and API. Our first flagship model with native multimodal support. Faster. More capable. Lower price. Open weights as usual.
-
@deepseek_ai
@deepseek_ai
on x
🧠 Asymmetric architecture. More intelligence, less cost. 🔹 552B-parameter MoE. 🔹 New Causal Encoder-Decoder architecture: just 8B active parameters for input, 16B for output. 🔹 New pre-training methods + larger-scale RL post-training deliver benchmark results ahead of flagship mo…
-
@zephyr_z9
@zephyr_z9
on x
The whale strikes back!!!
-
@deryatr_
Derya Unutmaz
on x
A significant part of the psy-op for the doomsday fearmongering is to stop this!👇But they won't succeed! I'm so happy DeepSeek just switched on the lights on this dark day. V4.1 Flash looks incredible! Better than even GLM-5.3, which is already an amazing open-source model!
-
@eriskiiii
Eris
on x
holy fucking shit this thing is fast. 200 TPS from the official fucking endpoint.
-
@ivanfioravanti
Ivan Fioravanti
on x
Unbelievable intelligence compressed in a Flash model! Probably going big big big is not the right answer. Smaller, less expensive and more capable models is the real way.
-
@miaai_lab
Mia
on x
DeepSeek v4.1 Flash is the smallest model in their new architecture. It weights 510 GB total. Which means 342 GB larger, or 204% bigger than DeepSeek v4 Flash Vision Exp.
-
@jun_song
Jun Song
on x
open-weight Flash model just mogged Opus-5 and GPT-5.6-Sol. This is insane.
-
@deepseek_ai
@deepseek_ai
on x
🌐 Supporting open source. Expanding deployment options. We'll work closely with the open-source community on V4.1-Flash inference support and explore more deployment options. Planning a large-scale deployment with 2,000 GPUs + a storage cluster? Let's talk. 🔹 Model: https://huggi…
-
@chrisgpt
Chris
on x
DeepSeek V4.1 Flash just dropped! It's a 552B MoE model, but activates only 8B parameters for input and 16B for output. And it's beating models like GPT-5.6 Sol and Opus 5 on several agentic/coding benchmarks. CyberGym: 88.1 vs 84.5 GPT-5.6 Sol Automation-Bench: 54.8 vs 45.8 Deep…
-
@ns123abc
Nik
on x
Deepseek just dropped v4.1 flash, fully open weights it beats gpt 5.6 sol, opus 5 and every chinese model on coding and cybersecurity at ~86x cheaper cost per million tokens running at 420-507 tok/s = faster than gemini 3.8 flash “smallest model in our new architecture family” bt…
-
@thom_wolf
Thomas Wolf
on x
don't get distracted by all the hedging words in its name ("flash", minor version): seems like DeepSeek V4.1 Flash is a major update weights at: https://huggingface.co/... paper at: https://huggingface.co/...
-
@theahmadosman
Ahmad
on x
This new DeepSeek V4.1 Flash model will make GPUs prices go even higher ... way higher $40k RTX PRO 6000 incoming Bookmark this tweet
-
@whoareme33
Nick Mykhailyshyn
on x
deepseek-v4.1-flash is insane! it found a 0day RCE in handlebars.js v4.7.9 in minutes for just $0.05 what's interesting is that deepseek-v4-pro-0813 needed multiple runs with the same prompt. flash found it consistently. security researchers, we're so cooked :)
-
@sheriyuo
Xiuyu Li
on x
DeepSeek V4.1 Flash, with its almost unreal RL results, has brought us back to an RL era that felt like it had disappeared. RL, rise again. BERT, rise again. https://huggingface.co/...
-
@testingcatalog
@testingcatalog
on x
DeepSeek V4.1 Flash is now available on Huggingface! > 552B parameters MoE model with a new Encoder-Decoder structure. > V4.1 Flash adopted a new pre-training method and underwent larger-scale reinforcement learning post-training. > The smallest model in DeepSeek new architecture…
-
@ypwang61
Yiping Wang
on x
A future where AGI is accessible and affordable enough to be used correctly by anyone, anywhere.
-
@zainhas
Zain
on x
oh wow have not seen this before for any model... this complicates things > “DeepSeek-V4.1-Flash supports a continuously controllable reasoning effort from 1 to 100.” reasoning_effort = [1 to 100]
-
r/LocalLLaMA
r
on reddit
deepseek-ai/DeepSeek-V4.1-Flash · Hugging Face
-
r/LocalLLM
r
on reddit
DeepSeek-V4.1-Flash is out
-
Deedy Das
Deedy Das
on linkedin
DeepSeek might be the new king of open-source models. It obliterates GLM 5.3 and Kimi K3 on benchmarks while being 4x and 10x cheaper. …
-
@bindureddy
Bindu Reddy
on x
Absolutely Insane! DeepSeek V4.1 Flash has astounding benchmark numbers It powers our personal agents and is close to FREE 😲 Get AI to write email, research topics, and pretty much do all your work and pay nothing
-
@jenzhuscott
Jen Zhu
on x
WTF 😅🚀😳 “98% of Astra's score at 1.4% of cost” Cancel the fucking IPOs 😂
-
@jessefelder.com
Jesse Felder
on bluesky
‘DeepSeek rolled out an AI model that charges as little as a fraction of a cent per million tokens, ramping up the pressure on rivals from Anthropic PBC to Z.AI Co.’ www.bloomberg.com/news/article... [image]