DeepSeek releases DeepSeek-V3, an open-source MoE model of 671B total parameters, with 37B activated per token, claiming it outperforms top models like GPT-4o
Chinese AI startup DeepSeek, known for challenging leading AI vendors with its innovative open-source technologies, today released a new ultra-large model: DeepSeek-V3.
VentureBeat Shubham Sharma
Context & Ripple Effects
This launch established DeepSeek-V3 as the base of a continuing release line: a later MIT-licensed V3 update clarified the model's open distribution, while later V3.2 work emphasized sparse attention and lower tool pricing.
The coverage also shows DeepSeek extending the same mixture-of-experts scale into a math-focused Prover-V2 release, rather than treating V3 as a one-off model.
First-order effects
- Developers can evaluate and deploy an open-source, sparse-activation model at a scale usually associated with leading closed-model providers; DeepSeek gains a public performance and efficiency benchmark, though its GPT-4o comparison remains the company's claim.
- The release makes DeepSeek-V3 a focal point for downstream fine-tuning, hosting, and benchmark testing, with its 37B-per-token activation figure central to cost comparisons.
Second-order effects
- Closed-model vendors and open-model rivals face more pressure to demonstrate performance-per-serving-cost, not just total parameter counts; buyers gain another model to include in procurement tests.
- Tooling and infrastructure providers can target MoE deployment and optimization as DeepSeek's subsequent sparse-attention V3.2 experiment shows the company continuing to compete on inference efficiency and price.
Third-order effects
- If comparable open-weight models keep improving, model capability may become less scarce at the base layer, shifting more differentiation toward distribution, specialized applications, and operational efficiency.
- The repeated V3-family releases point to an emerging competitive loop in which open licensing, architecture improvements, and lower serving costs reinforce one another; the durability of that shift depends on independent performance and deployment results.
The trend: This is an early data point in the shift from headline parameter scale toward open-weight models that compete through sparse architectures and lower-cost access.
Related: Open-weight complement economy · Model buyer power · DeepSeek · DeepSeek-V3 · DeepSeek releases MIT-licensed DeepSeek-V3-0324 · DeepSeek releases DeepSeek-V3.2-Exp
Related Coverage
- 1. Introduction — We present DeepSeek-V3, a strong Mixture-of-Experts (MoE) … GitHub
- Chinese AI company says breakthroughs enabled creating a leading-edge AI model with 11X less compute — DeepSeek's optimizations could highlight limits of US sanctions Tom's Hardware · Anton Shilov
- U.S' AI Hardware Restrictions on China Have Backfired Analytics India Magazine · Supreeth Koundinya
- Chinese start-up DeepSeek launches AI model that outperforms Meta, OpenAI products South China Morning Post · Ben Jiang
- DeepSeek's new AI model appears to be one of the best ‘open’ challengers yet TechCrunch · Kyle Wiggers
- DeepSeek-AI Just Released DeepSeek-V3: A Strong Mixture-of-Experts (MoE) Language Model with 671B Total Parameters with 37B Activated for Each Token MarkTechPost · Asif Razzaq
- DeepSeek-V3 sets new standard for open-source language models Neowin · Pradeep Viswanathan
- DeepSeek V3: New Open AI Model Surpasses Rivals and Challenges GPT-4o WinBuzzer · Markus Kasanmascheff
- DeepSeek-V3 Open-Source AI Model With Mixture-of-Experts Architecture Released Gadgets 360
- DeepSeek launches new AI model with 671 billion parameters, rivaling GPT-4o TechNode
- Debbie Downer Says, No AI Payoff Until 2026 Beyond Search · Stephen E. Arnold
- DeepSeek open-sources new AI model with 671B parameters SiliconANGLE · Maria Deutscher
Discussion
-
@digthatdata
David Marx
on bluesky
DeepSeek v3 released - model, code, and writeup github.com/deepseek-ai/... Some highlights: — * Multi-token prediction training objective — * FP8 + MoE for distributed training efficiency — * New “DualPipe” PP algorithm — * DeepSeekMoE + MH-Latent-A (see also DSv2) …
-
@karpathy
Andrej Karpathy
on x
DeepSeek (Chinese AI co) making it look easy today with an open weights release of a frontier-grade LLM trained on a joke of a budget (2048 GPUs for 2 months, $6M). For reference, this level of capability is supposed to require clusters of closer to 16K GPUs, the ones being brou…
-
@deepseek_ai
@deepseek_ai
on x
🚀 Introducing DeepSeek-V3! Biggest leap forward yet: ⚡ 60 tokens/second (3x faster than V2!) 💪 Enhanced capabilities 🛠 API compatibility intact 🌍 Fully open-source models & papers 🐋 1/n [image]
-
@emostaque
Emad
on x
To run DeepSeek v3 24/7 at 60 tokens per second (5x human reading speed) is $2 a day Latte ☕️ or AI 🤖 anon?
-
@tunguz
Bojan Tunguz
on x
All the export bans on high end semiconductors might have actually been counterproductive in the “worst” way imaginable. They seem to have forced Chinese researchers to be far more ingenious and resource efficient than they might have otherwise been. It also seems to confirm my
-
@amasad
Amjad Masad
on x
Biden's Chips Act constrained the performance of chips exported to China, so the Chinese innovated a way to train large models for cheap. Regulators never consider second order effects.
-
@jbohnslav
Jim Bohnslav
on x
What would the deepseek team accomplish with a month or two on xAI's H100 cluster?
-
@hosseeb
Haseeb
on x
Wow. Insanely good coding model, fully open source with only 37B active parameters. Beats Claude and GPT-4o on most benchmarks. China + open source is catching up... 2025 will be a crazy year.
-
@kimmonismus
@kimmonismus
on x
DeepSeek-V3! 60 tokens/second (3x faster than V2!) API compatibility intact Fully open-source models & papers 671B MoE parameters 37B activated parameters Trained on 14.8T high-quality tokens Beats Llama 3.1 405b on almost every benchmark [image]
-
@nrehiew_
@nrehiew_
on x
Arch wise they differ significantly from meta which just used a single massive dense transformer For oss Mixture of Experts, mixtral was the first (i think) and DeepSeek popularised it. Multi-Head Latent attention (MLA) comes from their Deepseek v2 paper which basically makes [im…
-
@nrehiew_
@nrehiew_
on x
This is the image that has been going around so you probably know how nuts this is but some added context is that Llama 3 405B was trained on 16K H100 https://x.com/... [image]
-
@mathemagic1an
Jay Hack
on x
DeepSeek v3 is an order of magnitude cheaper because it likely trained on frontier model outputs, in obvious violation of ToS ToS laundering by training on DeepSeek outputs is impossible to prevent. Does not bode well for economics of training frontier models
-
@drjimfan
@drjimfan
on x
Resource constraints are a beautiful thing. Survival instinct in a cut-throat AI competitive land is a prime drive for breakthroughs. I've been following DeepSeek for a long time. They had one of the best open coding models last year. Superior OSS models put huge pressure on
-
@lmsysorg
@lmsysorg
on x
The best open-source LLM, DeepSeek V3, has just been released! SGLang v0.4.1 is the officially recommended inference solution for it. The SGLang and DeepSeek teams worked together to support DeepSeek V3 FP8 on NVIDIA and AMD GPUs from day one. SGLang has supported MLA and DP
-
@charlesrollet1
Charles Rollet
on x
DeepSeek doesn't just pretend the Tiananmen Square massacre never happened. It also ‘firmly supports’ imprisoning prominent Uyghur dissidents for life. “Ilham Tohti is a person of significant notoriety in China... We firmly support the government's actions” [image]
-
@dnystedt
Dan Nystedt
on x
@tunguz Good question. CCP (Communist) propaganda always seeks to appear strong, no matter how weak. So are Deepseek's claims true? Is the AI robust? It could be all lies. But if not, that's bad news. https://x.com/...
-
@emollick
Ethan Mollick
on x
Unless things change dramatically, or there is some special sauce that is required & which the labs keep secret, frontier AI capabilities are likely to be available through open models (with the implication that no guardrails will hold in open models) all the way to possible AGI
-
@abhishekn
Abhishek Nagaraj
on x
we cover this in our paper on competition in generative AI ( https://www.nber.org/...) our key prediction was that moats cannot rely on “secret sauce” but must instead be predicated on complementary assets (compute, network effects ...)
-
@arankomatsuzaki
Aran Komatsuzaki
on x
Deepseek-V3-Base was just opensourced! - 685B MoE w/ 256 experts topk=8 with sigmoid routing - Outperforms Sonnet 3.5 on Aider benchmark https://huggingface.co/... [image]
-
@richardsocher
Richard Socher
on x
The race to lower LLM token prices continues with this impressive model trained on a much smaller budget. The higher the data quality, the lower the compute cost to train.
-
@amasad
Amjad Masad
on x
Summary of how DeepSeek V3 was so efficient at training a frontier-level model according to Perplexity: — DeepSeek-V3 was able to train their large 671B parameter model with a relatively low compute budget of 2.788M H800 GPU hours through several key innovations and
-
@kimmonismus
@kimmonismus
on x
Price-eval-ratio is next level with DeepSeek v3 One should not underestimate the importance of a good price for the LLMs so that they are really available to everyone and so that the models are widely accepted. [image]
-
@eladgil
Elad Gil
on x
Good evidence that a lot of efficiency being left on the table in US labs with massively scaled clusters
-
@theblackauroras
@theblackauroras
on x
It's not outperforming just open source models. It's o1 level and outperforms Claude.
-
@nearcyan
Near
on x
DeepSeek built this in an export-restricted cave with a pile of h800s [image]
-
@garybasin
Gary Basin
on x
The best thing about deepseek v3 is we'll now get sonnet 4 next week
-
@seunghyunseo7
Seunghyun Seo
on x
A quick summary of deepseek-v3 model (there may be wrong or missing details as this is based on a quick skim) tons of respect to their engineering team... [image]
-
@menhguin
Minh Nhat Nguyen
on x
at some point unis should just have deepseek technical papers as readings for ML/CS. it's hard NOT to gain alpha from that, let alone something so up-to-date and relevant to the frontier.
-
@balajis
Balaji
on x
In other words: the Chinese Deepseek paper references GPT-4o, which was in fact led by an Indian immigrant to the US. Many valid policy decisions one can make, but be real about the tradeoffs. 96% of the world is non-American and there is real talent out there. Original post:
-
@rasbt
Sebastian Raschka
on x
An updated back-of-the-envelope calculation of LLM pretraining costs based on the just-released DeepSeek-v3 report. And that doesn't even account for hyperparameter tuning, failed runs, or personnel costs. It really makes me appreciate the value of openly shared model weights! [i…
-
@casper_hansen_
Casper Hansen
on x
I'm DeepSeek-pilled now. Boggles my mind that Huggingface has not provided support for any DeepSeek version yet! [image]
-
@scaling01
@scaling01
on x
META could have trained DeepSeek-V3 at least 15 times using the compute budget of the Llama 3 model family ( 39.3 million H100 hours ) Meanwhile DeepSeek only spent 2.6 million H800 hours (a handicapped / worse H100) for a much better model [image]
-
@casper_hansen_
Casper Hansen
on x
Feels like Mistral could have dropped a model just as good as DeepSeek V3 but chose to develop products for revenue instead of focusing on research
-
@teknium1
@teknium1
on x
Looks like deepseek will make intelligence too cheap to meter instead of openai?
-
@terryyuezhuo
Terry Yue Zhuo
on x
Big congrats to @deepseek_ai! The V3 Chat model now ranks 1st on BigCodeBench-Hard. Complete — 40.5% Instruct — 28.4% Average — 34.5% Gemini-Exp-1206 Average — 34.1% o1-2024-12-17 (reasoning=medium) Average — 32.8% More results can be found at https://huggingface.co/... [image]
-
@reach_vb
@reach_vb
on x
LiveBench reported by r/LocalLlama - DeepSeek v3 is the BEST open weight LLM AND SECOND BEST non-reasoning LLM after ‘gemini-exp-1206’ 🔥 [image]
-
@tensor_fusion
Milton
on x
Peak engineering efficiency from the Whale. Napkin math: > DeepSeek-V3: 2048 H800s / 180K GPU-hours per trillion tokens > Llama 3: 16000 H100s for 54 days so ~1.3M GPU-hours per trillion tokens ~7.5x raw GPU efficiency in DeepSeek's favor. “and now we mog them”. [image]
-
@goodside
Riley Goodside
on x
This is such a vibes-based eval, but the first prompt I give any new LLM is “Which version is this?” and DeepSeek-V3 nailed it See below for how Claude, Gemini, ChatGPT, and Grok fare on the same — TLDR: it's all over the map [image]
-
@teortaxestex
@teortaxestex
on x
> $5.5M for Sonnet tier it's unsurprising that they're proud of it, but it sure feels like they're rubbing it in. «$100M runs, huh? 30.84M H100-hours on 405B, yeah? Half-witted Western hacks, your silicon is wasted on you, your thoughts wouldn't reduce loss of your own models» [i…
-
@balajis
Balaji
on x
China's Deepseek claims their new open source model was trained for just $5.6M, and that it's on par with GPT 4o and Claude 3.5 Sonnet. If true that's a >10X cost reduction.
-
@alexocheema
Alex Cheema
on x
I will run Deepseek-V3-Base 685B on M4 Mac Minis or die trying. 685B MoE with 256 experts — perfect for Apple Silicon since they have a lot of GPU memory and only a small subset of params are active at once. Outperforms SOTA closed-source models on benchmarks including Aider. [im…
-
@shravvmehtaa
Shrav Mehta
on x
Let's keep debating visas while China eats our lunch 🤡 China's leading LLM lab just dropped DeepSeek-V3. It's on par with GPT-4o and Sonnet 3.5 and cost < $10m to train.
-
@antimatter15
Kevin Kwok
on x
“DeepSeek was able to build this in a cave! With a box of scraps!”
-
@wordgrammer
@wordgrammer
on x
DeepSeek's new model seems to have proven this For the majority of startups. You cannot build a better datacenter than Microsoft. You cannot fundraise more than Sam or Elon. You cannot get more data than Google. But you actually can write 10x better code than any of them
-
@gallabytes
@gallabytes
on x
whale bros have the mandate of heaven, truly. xAI should just run the DeepSeek code on their cluster on the biggest dataset they can cobble together and release the best omni-model the world has ever seen. don't bother post-training it, just make the best base model and release
-
@simeon_cps
Siméon
on x
It's noteworthy that DeepSeek is the first competitive lab which is not created by former GoogleBrain/GDMs or OpenAI (and derivatives) staff.
-
@vincentweisser
Vincent Weisser
on x
Deepseek completely mogging the closed ai labs. Creating a state of the art model for $5.6m.
-
@reach_vb
@reach_vb
on x
The license allows commercial usage *without* an revenue caps, only asks one to respect the Use Policy! This makes DeepSeek V3 even more liberal than Llama 3x series 🔥 [image]
-
@jiayq
Yangqing Jia
on x
In 2019 I had a chat with the DeepSeek team, in the hope of selling them an AI cloud solution. I was trying to convince them a few things: - you don't need complicated cloud virtualization, you just need containers and an efficient scheduler. - you will need really fast,
-
@vin_sachi
Vin Sachidananda
on x
Interesting point on training costs/compute efficiency: Deepseek trained on $5.5M w sparse (MoE) architecture - only 38B active params Cheaper than llama 3.1 (dense) w better performance From Jeff Dean's NeurIPS talk, Gemini taking similar approach Maybe end of dense models
-
@hsu_steve
Steve Hsu
on x
DeepSeek V3 is out. Big improvements. This model is fast, inexpensive to run, and seems to beat Claude, 4o etc. on a broad set of benchmarks. Perhaps most importantly V3 is still open source! Keep in mind DS is a hedge fund whose founder (educated in AI at Zhejiang University)
-
@tom_doerr
Tom Dörr
on x
Deepseek V3 is amazing. It changed exactly what I wanted, thought it would take a lot more prompting to explain why I want to generate the answer first. Didn't even need to explain that it needs to be a separate signature [image]
-
@saranormous
@saranormous
on x
I don't think the US chip export controls are having their intended effect. Chinese model DeepSeek v3 very strong, and trained with OOM less money: “DeepSeek-V3 requires only 2.788M H800 GPU hours for its full training” (h800 is h100 with lower interchip bandwidth) [image]
-
@deanwball
Dean W. Ball
on x
The Chinese AGI lab DeepSeek is reporting an insanely low training cost of only $5.5 million for their new v3 model, which seems to match Claude 3.5 sonnet performance. DeepSeek also has a credible “o1-like” model. Again I will emphasize: if your policy solutions are [image]
-
@scaling01
@scaling01
on x
Some notes on the DeepSeek-V3 Technical Report :) The most insane thing to me: The whole training only cost $5.576 million or ~55 days on a 2048xH800 cluster. This is TINY compared to the Llama, GPT or Claude training runs. - 671B MoE with 37B activate params - DeepSeek MoE [imag…
-
@yigitkonur
@yigitkonur
on x
Bored during the holidays? For AI enthusiasts, it's the perfect time to catch up on the latest buzz 🤖 China's answer to OpenAI, Deepseek, just released their new v3 model, beating GPT-4o in most benchmarks It's now live on @ThinkbuddyAI for all curious minds 👇
-
@jsevillamol
Jaime Sevilla
on x
DeepSeek-V3 is impressively efficient. It has 37B active params and was pre-trained on 14.8T tokens, amounting to 6 x 37B x 14.8T = 3e24 FLOP. That's 10x less compute than llama 3.1 430B, yet better benchmark performance. [image]
-
@nikunj
Nikunj Kothari
on x
Feels like a Christmas miracle.. @deepseek_ai v3 - frontier OS model from China without throwing a LOT of compute at the problem “2.664M H800 GPU hours, we complete the pre-training of DeepSeek-V3 on 14.8T tokens, producing the currently strongest open-source base model.” [image]
-
@openrouterai
@openrouterai
on x
Deepseek has tripled in usage on OpenRouter since the v3 launch yesterday. Try it yourself, w/o subscription, including web search: [video]
-
@scaling01
@scaling01
on x
the real story here is that Grok-3 will use 100x the compute and just barely beat DeepSeek-V3 LMAO
-
@tom_doerr
Tom Dörr
on x
It's creepy how well Deepseek V3 understands what's going on without me having to explain anything. It suddenly feels like there's a ghost in the machine [image]
-
@dtrinh
Danny Trinh
on x
We're debating global skilled labor with a racist mob of locally unskilled laborers — meanwhile China put out Deepseek V3 with chewing gum and some paper clips. The US must win the global talent war and policies that get in the way only help our competition.
-
@anjneymidha
Anjney Midha
on x
Deepseek v3 seems to be a genuinely good model. China has caught up. Open source has caught up. The frontier costs abt $6M. I expect the model to crush it on @OpenRouterAI over the next few days. Lots of priors should be updated.
-
@jacobrintamaki
Jacob Rintamaki
on x
I think in many ways deepseek-v3 is a bigger update than o3 for “I didn't think that was possible yet.” There's so many great techniques they used as well; going to read their paper at least 4 times today. [image]
-
@artificialguybr
@artificialguybr
on x
How crazy is it to think that Deepseek is offering Deepseek V3 for 0.28/1M output and that this model is theoretically the second best non-reasoning model (second only to Gemini 1206) in the benchmarks. [image]
-
@casper_hansen_
Casper Hansen
on x
Deepseek just mogged all of us with one release. Where is Llama 4? Will it even be competitive?
-
@eturner303
Elliot Turner
on x
DeepSeek V3 was post-trained on synthetic CoT data generated by their R1 reasoning model. Here's some interesting tidbits on their data pipeline from the DSV3 paper [image]
-
@lm_zheng
Lianmin Zheng
on x
Highly respected! It is so impressive given the results and the very limited resources they have compared to other big labs. “DeepSeek-V3 is trained on a cluster equipped with 2048 NVIDIA H800 GPUs.”
-
@xlr8harder
@xlr8harder
on x
While operating under substantial hardware limitations in China, DeepSeek continues to execute flawlessly and remain consistently generous, sharing their crown jewels with the world. Incredible.
-
@nrehiew_
@nrehiew_
on x
How to train a 670B parameter model. Let's talk about the DeepSeek v3 report + some comparisons with what Meta did with Llama 405B [image]
-
@test_tm7873
@test_tm7873
on x
DeepSeek V3 is also one of the cheapest models ? Holyyyyyyy cat! [image]
-
@andrewcurran_
Andrew Curran
on x
The official release of the Whale, Deepseek V3 has arrived. [image]
-
@deedydas
Deedy
on x
The ex-quants in China dropped the best open-source LLM in the world: DeepSeek V3! — 671B params, 37B active MoE — On par with 3.5 Sonnet and 4o — $0.27/Mtok input, $1.1/Mtok output — 60tok/s — 128k context — Trained with just $5.5M on 14.8T tok! 53-page technical paper is GOLD […
-
@amasad
Amjad Masad
on x
Craziest thing is it took only $5.5m to train. US labs spend one — maybe two — order of magnitude more for frontier models.
-
@alexandr_wang
Alexandr Wang
on x
It is quite fitting that DeepSeek, China's leading LLM lab, releases its latest model V3 on Christmas. - on-par with GPT-4o & Claude 3.5 Sonnet - trained w/10x less compute The bitter lesson of Chinese tech: they work while America rests, and catch up cheaper, faster & stronger […
-
@tydsh
Yuandong Tian
on x
FP8 pre-training, MoE, strong performance with very limited budget, Distill from CoT for bootstrapping ... wow, this is great work 👏👏 👍👍
-
@itspaulai
Paul Couvert
on x
Wait, so we now have a 100% open source model that's better than GPT-4o?! DeepSeek v3 is even superior to Claude Sonnet 3.5 for code according to multiple benchmarks. Already available for free to everyone 🧵 [image]
-
@tim_dettmers
Tim Dettmers
on x
Reading the report, this is such clean engineering under resource constraints. The DeepSeek team directly engineered solutions to known problems under hardware constraints. All of this looks so elegant — no fancy “academic” solutions, just pure, solid engineering. Respect 👏