/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

DeepSeek releases DeepSeek-V3, an open-source MoE model of 671B total parameters, with 37B activated per token, claiming it outperforms top models like GPT-4o

Chinese AI startup DeepSeek, known for challenging leading AI vendors with its innovative open-source technologies, today released a new ultra-large model: DeepSeek-V3.

VentureBeat Shubham Sharma

Context & Ripple Effects

This launch established DeepSeek-V3 as the base of a continuing release line: a later MIT-licensed V3 update clarified the model's open distribution, while later V3.2 work emphasized sparse attention and lower tool pricing.

The coverage also shows DeepSeek extending the same mixture-of-experts scale into a math-focused Prover-V2 release, rather than treating V3 as a one-off model.

First-order effects

  • Developers can evaluate and deploy an open-source, sparse-activation model at a scale usually associated with leading closed-model providers; DeepSeek gains a public performance and efficiency benchmark, though its GPT-4o comparison remains the company's claim.
  • The release makes DeepSeek-V3 a focal point for downstream fine-tuning, hosting, and benchmark testing, with its 37B-per-token activation figure central to cost comparisons.

Second-order effects

  • Closed-model vendors and open-model rivals face more pressure to demonstrate performance-per-serving-cost, not just total parameter counts; buyers gain another model to include in procurement tests.
  • Tooling and infrastructure providers can target MoE deployment and optimization as DeepSeek's subsequent sparse-attention V3.2 experiment shows the company continuing to compete on inference efficiency and price.

Third-order effects

  • If comparable open-weight models keep improving, model capability may become less scarce at the base layer, shifting more differentiation toward distribution, specialized applications, and operational efficiency.
  • The repeated V3-family releases point to an emerging competitive loop in which open licensing, architecture improvements, and lower serving costs reinforce one another; the durability of that shift depends on independent performance and deployment results.

The trend: This is an early data point in the shift from headline parameter scale toward open-weight models that compete through sparse architectures and lower-cost access.

Discussion

  • @digthatdata David Marx on bluesky
    DeepSeek v3 released - model, code, and writeup github.com/deepseek-ai/...  Some highlights:  —  * Multi-token prediction training objective  —  * FP8 + MoE for distributed training efficiency  —  * New “DualPipe” PP algorithm  —  * DeepSeekMoE + MH-Latent-A (see also DSv2) …
  • @karpathy Andrej Karpathy on x
    DeepSeek (Chinese AI co) making it look easy today with an open weights release of a frontier-grade LLM trained on a joke of a budget (2048 GPUs for 2 months, $6M).  For reference, this level of capability is supposed to require clusters of closer to 16K GPUs, the ones being brou…
  • @deepseek_ai @deepseek_ai on x
    🚀 Introducing DeepSeek-V3! Biggest leap forward yet: ⚡ 60 tokens/second (3x faster than V2!) 💪 Enhanced capabilities 🛠 API compatibility intact 🌍 Fully open-source models & papers 🐋 1/n [image]
  • @emostaque Emad on x
    To run DeepSeek v3 24/7 at 60 tokens per second (5x human reading speed) is $2 a day Latte ☕️ or AI 🤖 anon?
  • @tunguz Bojan Tunguz on x
    All the export bans on high end semiconductors might have actually been counterproductive in the “worst” way imaginable. They seem to have forced Chinese researchers to be far more ingenious and resource efficient than they might have otherwise been. It also seems to confirm my
  • @amasad Amjad Masad on x
    Biden's Chips Act constrained the performance of chips exported to China, so the Chinese innovated a way to train large models for cheap. Regulators never consider second order effects.
  • @jbohnslav Jim Bohnslav on x
    What would the deepseek team accomplish with a month or two on xAI's H100 cluster?
  • @hosseeb Haseeb on x
    Wow. Insanely good coding model, fully open source with only 37B active parameters. Beats Claude and GPT-4o on most benchmarks. China + open source is catching up... 2025 will be a crazy year.
  • @kimmonismus @kimmonismus on x
    DeepSeek-V3! 60 tokens/second (3x faster than V2!) API compatibility intact Fully open-source models & papers 671B MoE parameters 37B activated parameters Trained on 14.8T high-quality tokens Beats Llama 3.1 405b on almost every benchmark [image]
  • @nrehiew_ @nrehiew_ on x
    Arch wise they differ significantly from meta which just used a single massive dense transformer For oss Mixture of Experts, mixtral was the first (i think) and DeepSeek popularised it. Multi-Head Latent attention (MLA) comes from their Deepseek v2 paper which basically makes [im…
  • @nrehiew_ @nrehiew_ on x
    This is the image that has been going around so you probably know how nuts this is but some added context is that Llama 3 405B was trained on 16K H100 https://x.com/... [image]
  • @mathemagic1an Jay Hack on x
    DeepSeek v3 is an order of magnitude cheaper because it likely trained on frontier model outputs, in obvious violation of ToS ToS laundering by training on DeepSeek outputs is impossible to prevent. Does not bode well for economics of training frontier models
  • @drjimfan @drjimfan on x
    Resource constraints are a beautiful thing. Survival instinct in a cut-throat AI competitive land is a prime drive for breakthroughs. I've been following DeepSeek for a long time. They had one of the best open coding models last year. Superior OSS models put huge pressure on
  • @lmsysorg @lmsysorg on x
    The best open-source LLM, DeepSeek V3, has just been released! SGLang v0.4.1 is the officially recommended inference solution for it. The SGLang and DeepSeek teams worked together to support DeepSeek V3 FP8 on NVIDIA and AMD GPUs from day one. SGLang has supported MLA and DP
  • @charlesrollet1 Charles Rollet on x
    DeepSeek doesn't just pretend the Tiananmen Square massacre never happened. It also ‘firmly supports’ imprisoning prominent Uyghur dissidents for life. “Ilham Tohti is a person of significant notoriety in China... We firmly support the government's actions” [image]
  • @dnystedt Dan Nystedt on x
    @tunguz Good question. CCP (Communist) propaganda always seeks to appear strong, no matter how weak. So are Deepseek's claims true? Is the AI robust? It could be all lies. But if not, that's bad news. https://x.com/...
  • @emollick Ethan Mollick on x
    Unless things change dramatically, or there is some special sauce that is required & which the labs keep secret, frontier AI capabilities are likely to be available through open models (with the implication that no guardrails will hold in open models) all the way to possible AGI
  • @abhishekn Abhishek Nagaraj on x
    we cover this in our paper on competition in generative AI ( https://www.nber.org/...) our key prediction was that moats cannot rely on “secret sauce” but must instead be predicated on complementary assets (compute, network effects ...)
  • @arankomatsuzaki Aran Komatsuzaki on x
    Deepseek-V3-Base was just opensourced! - 685B MoE w/ 256 experts topk=8 with sigmoid routing - Outperforms Sonnet 3.5 on Aider benchmark https://huggingface.co/... [image]
  • @richardsocher Richard Socher on x
    The race to lower LLM token prices continues with this impressive model trained on a much smaller budget. The higher the data quality, the lower the compute cost to train.
  • @amasad Amjad Masad on x
    Summary of how DeepSeek V3 was so efficient at training a frontier-level model according to Perplexity: — DeepSeek-V3 was able to train their large 671B parameter model with a relatively low compute budget of 2.788M H800 GPU hours through several key innovations and
  • @kimmonismus @kimmonismus on x
    Price-eval-ratio is next level with DeepSeek v3 One should not underestimate the importance of a good price for the LLMs so that they are really available to everyone and so that the models are widely accepted. [image]
  • @eladgil Elad Gil on x
    Good evidence that a lot of efficiency being left on the table in US labs with massively scaled clusters
  • @theblackauroras @theblackauroras on x
    It's not outperforming just open source models. It's o1 level and outperforms Claude.
  • @nearcyan Near on x
    DeepSeek built this in an export-restricted cave with a pile of h800s [image]
  • @garybasin Gary Basin on x
    The best thing about deepseek v3 is we'll now get sonnet 4 next week
  • @seunghyunseo7 Seunghyun Seo on x
    A quick summary of deepseek-v3 model (there may be wrong or missing details as this is based on a quick skim) tons of respect to their engineering team... [image]
  • @menhguin Minh Nhat Nguyen on x
    at some point unis should just have deepseek technical papers as readings for ML/CS. it's hard NOT to gain alpha from that, let alone something so up-to-date and relevant to the frontier.
  • @balajis Balaji on x
    In other words: the Chinese Deepseek paper references GPT-4o, which was in fact led by an Indian immigrant to the US. Many valid policy decisions one can make, but be real about the tradeoffs. 96% of the world is non-American and there is real talent out there. Original post:
  • @rasbt Sebastian Raschka on x
    An updated back-of-the-envelope calculation of LLM pretraining costs based on the just-released DeepSeek-v3 report. And that doesn't even account for hyperparameter tuning, failed runs, or personnel costs. It really makes me appreciate the value of openly shared model weights! [i…
  • @casper_hansen_ Casper Hansen on x
    I'm DeepSeek-pilled now. Boggles my mind that Huggingface has not provided support for any DeepSeek version yet! [image]
  • @scaling01 @scaling01 on x
    META could have trained DeepSeek-V3 at least 15 times using the compute budget of the Llama 3 model family ( 39.3 million H100 hours ) Meanwhile DeepSeek only spent 2.6 million H800 hours (a handicapped / worse H100) for a much better model [image]
  • @casper_hansen_ Casper Hansen on x
    Feels like Mistral could have dropped a model just as good as DeepSeek V3 but chose to develop products for revenue instead of focusing on research
  • @teknium1 @teknium1 on x
    Looks like deepseek will make intelligence too cheap to meter instead of openai?
  • @terryyuezhuo Terry Yue Zhuo on x
    Big congrats to @deepseek_ai! The V3 Chat model now ranks 1st on BigCodeBench-Hard. Complete — 40.5% Instruct — 28.4% Average — 34.5% Gemini-Exp-1206 Average — 34.1% o1-2024-12-17 (reasoning=medium) Average — 32.8% More results can be found at https://huggingface.co/... [image]
  • @reach_vb @reach_vb on x
    LiveBench reported by r/LocalLlama - DeepSeek v3 is the BEST open weight LLM AND SECOND BEST non-reasoning LLM after ‘gemini-exp-1206’ 🔥 [image]
  • @tensor_fusion Milton on x
    Peak engineering efficiency from the Whale. Napkin math: > DeepSeek-V3: 2048 H800s / 180K GPU-hours per trillion tokens > Llama 3: 16000 H100s for 54 days so ~1.3M GPU-hours per trillion tokens ~7.5x raw GPU efficiency in DeepSeek's favor. “and now we mog them”. [image]
  • @goodside Riley Goodside on x
    This is such a vibes-based eval, but the first prompt I give any new LLM is “Which version is this?” and DeepSeek-V3 nailed it See below for how Claude, Gemini, ChatGPT, and Grok fare on the same — TLDR: it's all over the map [image]
  • @teortaxestex @teortaxestex on x
    > $5.5M for Sonnet tier it's unsurprising that they're proud of it, but it sure feels like they're rubbing it in. «$100M runs, huh? 30.84M H100-hours on 405B, yeah? Half-witted Western hacks, your silicon is wasted on you, your thoughts wouldn't reduce loss of your own models» [i…
  • @balajis Balaji on x
    China's Deepseek claims their new open source model was trained for just $5.6M, and that it's on par with GPT 4o and Claude 3.5 Sonnet. If true that's a >10X cost reduction.
  • @alexocheema Alex Cheema on x
    I will run Deepseek-V3-Base 685B on M4 Mac Minis or die trying. 685B MoE with 256 experts — perfect for Apple Silicon since they have a lot of GPU memory and only a small subset of params are active at once. Outperforms SOTA closed-source models on benchmarks including Aider. [im…
  • @shravvmehtaa Shrav Mehta on x
    Let's keep debating visas while China eats our lunch 🤡 China's leading LLM lab just dropped DeepSeek-V3. It's on par with GPT-4o and Sonnet 3.5 and cost < $10m to train.
  • @antimatter15 Kevin Kwok on x
    “DeepSeek was able to build this in a cave! With a box of scraps!”
  • @wordgrammer @wordgrammer on x
    DeepSeek's new model seems to have proven this For the majority of startups. You cannot build a better datacenter than Microsoft. You cannot fundraise more than Sam or Elon. You cannot get more data than Google. But you actually can write 10x better code than any of them
  • @gallabytes @gallabytes on x
    whale bros have the mandate of heaven, truly. xAI should just run the DeepSeek code on their cluster on the biggest dataset they can cobble together and release the best omni-model the world has ever seen. don't bother post-training it, just make the best base model and release
  • @simeon_cps Siméon on x
    It's noteworthy that DeepSeek is the first competitive lab which is not created by former GoogleBrain/GDMs or OpenAI (and derivatives) staff.
  • @vincentweisser Vincent Weisser on x
    Deepseek completely mogging the closed ai labs. Creating a state of the art model for $5.6m.
  • @reach_vb @reach_vb on x
    The license allows commercial usage *without* an revenue caps, only asks one to respect the Use Policy! This makes DeepSeek V3 even more liberal than Llama 3x series 🔥 [image]
  • @jiayq Yangqing Jia on x
    In 2019 I had a chat with the DeepSeek team, in the hope of selling them an AI cloud solution. I was trying to convince them a few things: - you don't need complicated cloud virtualization, you just need containers and an efficient scheduler. - you will need really fast,
  • @vin_sachi Vin Sachidananda on x
    Interesting point on training costs/compute efficiency: Deepseek trained on $5.5M w sparse (MoE) architecture - only 38B active params Cheaper than llama 3.1 (dense) w better performance From Jeff Dean's NeurIPS talk, Gemini taking similar approach Maybe end of dense models
  • @hsu_steve Steve Hsu on x
    DeepSeek V3 is out. Big improvements. This model is fast, inexpensive to run, and seems to beat Claude, 4o etc. on a broad set of benchmarks. Perhaps most importantly V3 is still open source! Keep in mind DS is a hedge fund whose founder (educated in AI at Zhejiang University)
  • @tom_doerr Tom Dörr on x
    Deepseek V3 is amazing. It changed exactly what I wanted, thought it would take a lot more prompting to explain why I want to generate the answer first. Didn't even need to explain that it needs to be a separate signature [image]
  • @saranormous @saranormous on x
    I don't think the US chip export controls are having their intended effect. Chinese model DeepSeek v3 very strong, and trained with OOM less money: “DeepSeek-V3 requires only 2.788M H800 GPU hours for its full training” (h800 is h100 with lower interchip bandwidth) [image]
  • @deanwball Dean W. Ball on x
    The Chinese AGI lab DeepSeek is reporting an insanely low training cost of only $5.5 million for their new v3 model, which seems to match Claude 3.5 sonnet performance. DeepSeek also has a credible “o1-like” model. Again I will emphasize: if your policy solutions are [image]
  • @scaling01 @scaling01 on x
    Some notes on the DeepSeek-V3 Technical Report :) The most insane thing to me: The whole training only cost $5.576 million or ~55 days on a 2048xH800 cluster. This is TINY compared to the Llama, GPT or Claude training runs. - 671B MoE with 37B activate params - DeepSeek MoE [imag…
  • @yigitkonur @yigitkonur on x
    Bored during the holidays? For AI enthusiasts, it's the perfect time to catch up on the latest buzz 🤖 China's answer to OpenAI, Deepseek, just released their new v3 model, beating GPT-4o in most benchmarks It's now live on @ThinkbuddyAI for all curious minds 👇
  • @jsevillamol Jaime Sevilla on x
    DeepSeek-V3 is impressively efficient. It has 37B active params and was pre-trained on 14.8T tokens, amounting to 6 x 37B x 14.8T = 3e24 FLOP. That's 10x less compute than llama 3.1 430B, yet better benchmark performance. [image]
  • @nikunj Nikunj Kothari on x
    Feels like a Christmas miracle.. @deepseek_ai v3 - frontier OS model from China without throwing a LOT of compute at the problem “2.664M H800 GPU hours, we complete the pre-training of DeepSeek-V3 on 14.8T tokens, producing the currently strongest open-source base model.” [image]
  • @openrouterai @openrouterai on x
    Deepseek has tripled in usage on OpenRouter since the v3 launch yesterday. Try it yourself, w/o subscription, including web search: [video]
  • @scaling01 @scaling01 on x
    the real story here is that Grok-3 will use 100x the compute and just barely beat DeepSeek-V3 LMAO
  • @tom_doerr Tom Dörr on x
    It's creepy how well Deepseek V3 understands what's going on without me having to explain anything. It suddenly feels like there's a ghost in the machine [image]
  • @dtrinh Danny Trinh on x
    We're debating global skilled labor with a racist mob of locally unskilled laborers — meanwhile China put out Deepseek V3 with chewing gum and some paper clips. The US must win the global talent war and policies that get in the way only help our competition.
  • @anjneymidha Anjney Midha on x
    Deepseek v3 seems to be a genuinely good model. China has caught up. Open source has caught up. The frontier costs abt $6M. I expect the model to crush it on @OpenRouterAI over the next few days. Lots of priors should be updated.
  • @jacobrintamaki Jacob Rintamaki on x
    I think in many ways deepseek-v3 is a bigger update than o3 for “I didn't think that was possible yet.” There's so many great techniques they used as well; going to read their paper at least 4 times today. [image]
  • @artificialguybr @artificialguybr on x
    How crazy is it to think that Deepseek is offering Deepseek V3 for 0.28/1M output and that this model is theoretically the second best non-reasoning model (second only to Gemini 1206) in the benchmarks. [image]
  • @casper_hansen_ Casper Hansen on x
    Deepseek just mogged all of us with one release. Where is Llama 4? Will it even be competitive?
  • @eturner303 Elliot Turner on x
    DeepSeek V3 was post-trained on synthetic CoT data generated by their R1 reasoning model. Here's some interesting tidbits on their data pipeline from the DSV3 paper [image]
  • @lm_zheng Lianmin Zheng on x
    Highly respected! It is so impressive given the results and the very limited resources they have compared to other big labs. “DeepSeek-V3 is trained on a cluster equipped with 2048 NVIDIA H800 GPUs.”
  • @xlr8harder @xlr8harder on x
    While operating under substantial hardware limitations in China, DeepSeek continues to execute flawlessly and remain consistently generous, sharing their crown jewels with the world. Incredible.
  • @nrehiew_ @nrehiew_ on x
    How to train a 670B parameter model. Let's talk about the DeepSeek v3 report + some comparisons with what Meta did with Llama 405B [image]
  • @test_tm7873 @test_tm7873 on x
    DeepSeek V3 is also one of the cheapest models ? Holyyyyyyy cat! [image]
  • @andrewcurran_ Andrew Curran on x
    The official release of the Whale, Deepseek V3 has arrived. [image]
  • @deedydas Deedy on x
    The ex-quants in China dropped the best open-source LLM in the world: DeepSeek V3! — 671B params, 37B active MoE — On par with 3.5 Sonnet and 4o — $0.27/Mtok input, $1.1/Mtok output — 60tok/s — 128k context — Trained with just $5.5M on 14.8T tok! 53-page technical paper is GOLD […
  • @amasad Amjad Masad on x
    Craziest thing is it took only $5.5m to train. US labs spend one — maybe two — order of magnitude more for frontier models.
  • @alexandr_wang Alexandr Wang on x
    It is quite fitting that DeepSeek, China's leading LLM lab, releases its latest model V3 on Christmas. - on-par with GPT-4o & Claude 3.5 Sonnet - trained w/10x less compute The bitter lesson of Chinese tech: they work while America rests, and catch up cheaper, faster & stronger […
  • @tydsh Yuandong Tian on x
    FP8 pre-training, MoE, strong performance with very limited budget, Distill from CoT for bootstrapping ... wow, this is great work 👏👏 👍👍
  • @itspaulai Paul Couvert on x
    Wait, so we now have a 100% open source model that's better than GPT-4o?! DeepSeek v3 is even superior to Claude Sonnet 3.5 for code according to multiple benchmarks. Already available for free to everyone 🧵 [image]
  • @tim_dettmers Tim Dettmers on x
    Reading the report, this is such clean engineering under resource constraints. The DeepSeek team directly engineered solutions to known problems under hardware constraints. All of this looks so elegant — no fancy “academic” solutions, just pure, solid engineering. Respect 👏