/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

Rather than weakening China's AI capabilities, US sanctions appear to be driving startups like DeepSeek to innovate by prioritizing efficiency and collaboration

The AI community is abuzz over DeepSeek R1, a new open-source reasoning model.  —  The model was developed by the Chinese AI startup DeepSeek …

MIT Technology Review Caiwei Chen

Context & Ripple Effects

DeepSeek’s R1 arrives after reporting that export controls had already pushed the company to build DeepSeek-V3 without the newest chips, making efficiency under hardware constraints a central part of its development story.

The release also foreshadows a broader Chinese response: later coverage describes Alibaba, Baidu, and DeepSeek using open source to work around curbs and draw in outside refinement. That makes collaboration a strategic distribution choice, not merely a model-release format.

First-order effects

  • DeepSeek gains a visible open-source reasoning-model release through R1, while its engineering approach is framed around extracting more capability from constrained compute resources.
  • US restrictions impose a sharper incentive on Chinese AI startups to prioritize efficiency and collaborative development rather than relying on access to the latest chips.

Second-order effects

Third-order effects

  • If this pattern persists, export controls may shape the technical and distribution architecture of China’s AI sector—favoring compute-efficient models and broader model sharing—rather than simply determining whether capable models emerge.
  • The result could deepen a split between AI ecosystems organized around controlled access to advanced hardware and ones that seek resilience through open distribution, though the durability of that split depends on continued developer adoption and model performance.

The trend: AI export restrictions are increasingly acting as an innovation constraint that can redirect model development toward efficiency, open-source collaboration, and ecosystem resilience.

Discussion

  • @prietschka Paul Rietschka on bluesky
    The overinvestment in ML research + researchers in China has echoes in the extreme overinvestment in real estate in the country.  —  I mean, are Americans just now noting the prevalence of Chinese authors in the torrent of AI papers released daily?  —  This overinvestment won't e…
  • @willdouglasheaven Will Douglas Heaven on bluesky
    “The rapid evolution of AI demands agility from Chinese firms to survive.”  —  www.technologyreview.com/2025/01/24/ 1...  Good DeepSeek explainer from @caiwei.bsky.social
  • @RuthMalan@mastodon.social Ruth Malan on mastodon
    “The model was developed by the Chinese AI startup DeepSeek, which claims that R1 matches or even surpasses OpenAI's ChatGPT o1 on multiple key benchmarks but operates at a fraction of the cost.”  —  “To create R1, DeepSeek had to rework its training process to reduce the strain …
  • @naval @naval on x
    Turns out that instead of scraping the web to train an AI, you can just scrape the AI that scraped the web.
  • @armanddoma Armand Domalewski on x
    every single person on this list should get a visit from a US government agent holding a bag of cash in one hand and a bag of visas in the other [image]
  • @nickadobos Nick Dobos on x
    Deepseek is currently #3 in the AppStore (productivity section) Google Gemini is #5 Perplexity #36 Grok #37 Claude #44 Infer what you will about this information [image]
  • @benioff Marc Benioff on x
    Deepseek reshines the spotlight on the true treasure of AI: it's not the UI or the model—those are just commodities. The real value, the oxygen that gives AI life, lies in the data and the metadata that gives the model its context and power. Just like oxygen sustains us, data
  • @tphuang @tphuang on x
    Deepseek app is now in top 10 on iOS App Store & #3 among productivity apps Large part of DS future value will come from its app. Not making money now, but it can monetize that in the future like google. It has huge cost advantage over OpenAI. Can put latest model on its app. [im…
  • @jbulltard1 @jbulltard1 on x
    the fact that $META raised Capex guidance yesterday and closed green after all the deep seek nonsense everyone on fintwit is drooling over is enough to tell you what real money thinks of another Chinese gimmick. I remember when Temu was going to kill Amazon 12 months ago too.
  • @matthewstoller Matt Stoller on x
    Deepseek is forcing the entire Silicon Valley ecosystem to recognize that Lina Khan was right, even if they won't admit it.
  • @matthewstoller Matt Stoller on x
    American big tech firms are bad at building things because their focus is not on building things, it's on monopolization and political power. No different than Boeing. This has been obvious for years. https://www.thebignewsletter.com/ ...
  • @buccocapital @buccocapital on x
    First time laughing out loud at a community note [image]
  • @teroterotero Tero Kuittinen on x
    China undermined his half a trillion desert crystal castle and dude is now posting Napoleon quotes with implied vocal fry
  • @garrytan Garry Tan on x
    Do people really believe this? If training models get cheaper faster and easier, the demand for inference (actual real world use of AI) will grow and accelerate even faster, which assures the supply of compute will be used
  • @riyanmendonsa Riyan Mendonsa on x
    Looks like DeepSeek R1 had a wider knowledge base as well, or at least a preference for eastern philosophy. Gives it an interesting edge being so well rounded. Curious how its biases are too 🤔
  • @davidsholz David on x
    in my testing, deepseek crushes western models on ancient chinese philosophy and literature, while also having a much stronger command of english than my first-hand chinese sources. it feels like communing with literary/historical/philosophical knowledge across generations that i
  • @teknium1 @teknium1 on x
    Unbelievable the amount of cope, seethe, and hoop jumping people are doing to discredit deepseeks accomplishments lol
  • @balajis Balaji on x
    China has a 4000 year old civilization. They had a terrible 20th century thanks to Maoism, but they really are capable. You don't get to global #1 in manufacturing by accident. Even if (especially if!) you consider them a geopolitical rival, their engineers deserve respect.
  • @gonglei89 Lei Gong on x
    Tech bro culture in SV today is unrecognizable from the STEM nerd culture I grew up with. Until Americans figure out these are different things I'm skeptical they'll find quick answers against China. For decades they thought they were Tony Stark when they were actually Mysterio.
  • @cloneofsimo Simo Ryu on x
    As a foregner its really weird to see americans saying this is huge L for America or something If you read their paper, all methodologies really just came from google (Shazeer to be honest), openai and Meta (MoE, Transformer, MTP, scaling law, pytorch, even the improved NCCL,
  • @garymarcus Gary Marcus on x
    Interesting take - but I disagree. R1 is really impressive work, and quite possibly in some ways a game changer, but reasoning is NOT a solved problem. A really robust solution to reasoning would truly be a moat, but I don't think we have seen that yet.
  • @broseph_stalin Ashok Kumar on x
    For those confused why the US state and Silicon Valley are having a meltdown on twitter today. China released multiple AI models that are 50x more efficient than the best American AI models and made them open source, ruining the AI market
  • @ryangrim Ryan Grim on x
    This doesn't seem to be getting enough attention. The Silicon Valley social contract forced on the public by Obama and then Trump and then Biden (minus Lina Khan) and now Trump was straight forward: We will let these bros become the richest people in human history and in exchange
  • @deedydas Deedy on x
    America just can't imagine a world where China innovates
  • r/neoliberal r on reddit
    How a top Chinese AI model overcame US sanctions
  • @shibochentech Shibo Chen on x
    This is a dangerous thinking. Good and open sourced science breakthrough should be cherished, not viewed as geopolitical competition.
  • @jfpuget @jfpuget on x
    Every breakthrough in AI was in the US? Wasn't SGD a breakthrough? (from Leon Botou in France) Weren't CNN a breakthrough? (from Yann Lecun in France) Wasn't stable diffusion a breakthrough? (From Germany) Wasn't ViT a breathrough? (from Switzerland, at least partly) Wasn't
  • @teknium1 @teknium1 on x
    Has the USA considered accelerating OpenSource? How bout we do that. Biggest driver of innovation and cost reduction - AND - has the added benefit of equalizing access to AI to all - AND - will keep the research community working within and with the US? Or will we fall for
  • @basedjensen @basedjensen on x
    I am sorry export controls simply will not work against people who can pull this off. [image]
  • @alexandr_wang Alexandr Wang on x
    DeepSeek is a wake up call for America, but it doesn't change the strategy: - USA must out-innovate &race faster, as we have done in the entire history of AI - Tighten export controls on chips so that we can maintain future leads Every major breakthrough in AI has been American
  • @infanzone @infanzone on bluesky
    “When Chinese quant hedge fund founder Liang Wenfeng got into AI research, he took 10,000 Nvidia chips and assembled a team of young, ambitious talent.  Two years later, DeepSeek exploded onto the scene.”  [embedded post]
  • @edzitron.com Ed Zitron on bluesky
    To be fair they can barely find one for ChatGPT [embedded post]
  • @zoeschiffer Zoë Schiffer on bluesky
    “I wouldn't be able to find a commercial reason for founding DeepSeek even if you ask me to,” said founder Liang Wenfeng: www.wired.com/story/deepse...
  • @ErikJonker@mastodon.social Erik Jonker on mastodon
    Deepseek shows you can compete with Bigtech by being smarter.  Instead of only having more compute, data and money.  Which is a positive message also for AI startups in Europe.  Only too bad Deepseek is from China.  —  #ai #deepseek  —  https://www.wired.com/...
  • @deepsailcapital @deepsailcapital on x
    There are three options for what happened at DeepSeek: 1) They black ended an API from Llama or another open source LLM to help train their model, which means they just “borrowed” the LLM training logic. 2) They actually have a lot of H100s (as per Scale CEO has eluded too).
  • @omooretweets Olivia Moore on x
    DeepSeek's mobile app has entered the top 10 of the U.S. App Store. It's getting ~300k global daily downloads. This may be the first non-GPT based assistant to get mainstream U.S. usage. Claude has not cracked the top 200. [image]
  • r/LocalLLaMA r on reddit
    How Chinese AI Startup DeepSeek Made a Model that Rivals OpenAI
  • @icooper Ian Cooper on bluesky
    The story goes that British software devs in gaming (and SFX artists) became world beaters, allowing the UK to bat outside its league, because they had to practice their craft operating with far fewer resources than their US counterparts.  —  China is pulling the same trick.  —  …
  • @pmarca Marc Andreessen on x
    Deepseek R1 is one of the most amazing and impressive breakthroughs I've ever seen — and as open source, a profound gift to the world. 🤖🫡
  • @jjitsev @jjitsev on x
    (Yet) another tale of Rise and Fall: DeepSeek R1 is claimed to match o1/o1-preview on olympiad level math & coding problems. Can it handle versions of AIW problems that reveal generalization & basic reasoning deficits in SOTA LLMs? ( https://arxiv.org/...) 🧵1/n [image]
  • @nealkhosla Neal Khosla on x
    deepseek is a ccp state psyop + economic warfare to make american ai unprofitable they are faking the cost was low to justify setting price low and hoping everyone switches to it damage AI competitiveness in the us dont take the bait
  • @dorialexander Alexander Doria on x
    So DeepSeek situation summarized: *They are not a small engineer team but one of the leading frontier lab (+100 researchers full time). *They are not a newcomer. Started in 2023 by retraining a llama, then slowly rising to the top. All documented in their 16 (!) papers.
  • @garymarcus Gary Marcus on x
    Important, smart thread on DeepSeek R1 and generalization.
  • @kimmonismus @kimmonismus on x
    Billionaire and Scale AI CEO Alexandr Wang: DeepSeek has about 50,000 NVIDIA H100s that they can't talk about because of the US export controls that are in place. [video]
  • @dorialexander Alexander Doria on x
    Ah seeing multiple critics for stating that 2023 born DeepSeek is not a “newcomer”. I'm sorry if you are not aware of how fast theses things go, you're not really into LLM research.
  • @joshc0301 Josh on x
    Apologists for capitalism say that capitalism “incentivizes innovation” when research is best done collaboratively for the goal of advancing humanity When the primary goal for research is profit, the default is to avoid sharing info to competitors, which slows innovation
  • @dorialexander Alexander Doria on x
    The wave of R1 reproduction is just starting and impressive results already. This will be a massive boon for small models (even more so with some level of specialization).
  • @sivil_taram Qian Liu on x
    🚀 After 5 days of DeepSeek-R1, we've replicated its pure reinforcement learning magic on math reasoning — no reward models, no supervised fine-tuning, from a base model — and the results are mind-blowing: 🧠 A 7B model + 8K MATH examples for verification + Reinforcement
  • @suspendedrobot @suspendedrobot on x
    OpenAI stole from the whole internet to make itself richer, DeepSeek stole from them and give it back to the masses for free I think there is a certain british folktale about this
  • @philippilk Philip Pilkington on x
    The responses to this have been nuts. There are apparently a ton of AI hype-beasts on Twitter who have literally no idea what to do now that a few underfunded Chinese guys blew up their collective grift. Their response is now that DeepSeek is some sort of CCCP psyop. 🥴😵‍💫
  • @petergyang Peter Yang on x
    I find DeepSeek's thinking output more fascinating than its actual output
  • @itspaulai Paul Couvert on x
    No need to pay $200 to use Operator You can create an agent that uses a web browser without writing a line of code. Combine DeepSeek R1 and Browser Use (free and open source) and you're good to go. (Links and prompt below) [video]
  • @nielsrogge Niels Rogge on x
    “So there's this Chinese company called DeepSeek which basically does what OpenAI initially intended to do. They open-sourced a model trained with large-scale reinforcement learning, beating everyone else, and even releasing a paper detailing their process” [image]
  • @soumithchintala Soumith Chintala on x
    i'm comically impressed that people are coping on deepseek by spewing bizarre conspiracy theories — despite deepseek open-sourcing and writing some of the most detail oriented papers ever. read. replicate. compete. don't be salty, just makes you look incompetent.
  • @jaycaspiankang Kang on x
    This DeepSeek story is hilarious. Greatest troll job in years. They just tweeted out the secrets and now what's gonna happen to NVIDIA and open ai? I guess we can still use it to make funny college football memes.
  • @agnostoxxx Le Shrub on x
    Everyone is accusing Deepseek of either being a fraud or a CCP Trojan Horse. Meanwhile, Sam Altman is openly vying to become your new AI Overlord and y'all cool with it 🥹 [image]
  • @gregjstoker Greg J Stoker on x
    The meltdown on X by American techno-feudalists has to do with the Chinese AI model “Deepseek” being far more efficient and cheaper than anything coming out of Silicon Valley. Technical leadership v.s the American business-led model designed to absorb the maximum amount of [image…
  • @yishan @yishan on x
    R1 isn't the true big disruption. It's just a herald. Here's what I predict will happen: Within 3-6 months, an American company will duplicate the same cost-savings in a similar model. DS literally told everyone how to do it. After that, Deepseek will come out with something so
  • @archiexzzz Archie Sengupta on x
    DeepSeek Engineers [image]
  • @nathanbenaich Nathan Benaich on x
    hot take: deepseek is a blessing for ai startups that can now rip out their costly americano models and have profitable unit economics their investors should be happy, happy as a hippo [image]
  • @suhail @suhail on x
    Looks like DeepSeek just literally did it more efficiently. Game recognize game. [image]
  • @vrushankdes Vrushank Desai on x
    i find it hilarious that Deepseek released pages of details about their parallelism/quantization/etc to stave off haters doubting their training efficiency and.....haters still can't believe how efficient their model is lmaoo [image]
  • @amasad Amjad Masad on x
    So much cope about DeepSeek. Not only did they release a great model. they also released a breakthrough training method (R1 Zero) that's already reproducing. I doubt they lied about training costs, but even if they did they're still awesome for this great gift to the world.
  • @iamgingertrash @iamgingertrash on x
    Sam spent more on this than Deepseek did to train the model that killed OpenAI [image]
  • @iamgingertrash @iamgingertrash on x
    This guy is a product of nepotism and his dad is balls-deep in OpenAI stock Can you smell the cope? Here's his next tweet: “It's patriotic to use OpenAI and you're a communist if you use open source models like Deepseek”
  • @nxthompson @nxthompson on x
    The anger over Deepseek training on ChatGPT's outputs, in violation of TOS, is justified. But it might perhaps also make the AI companies think about the training they did on copyrighted content without consent. https://techcrunch.com/...
  • @natfriedman Nat Friedman on x
    The deepseek team is obviously really good. China is full of talented engineers. Every other take is cope. Sorry.
  • @hardmaru @hardmaru on x
    DeepSeek is a side project 🔥 [image]
  • @hxiao Han Xiao on x
    @abacaj deepseek's holding 幻方量化 is a quant company, many years already,super smart guys with top math background; happened to own a lot GPU for trading/mining purpose, and deepseek is their side project for squeezing those gpus
  • @michael_kove Michael Kove on x
    DeepSeek stole the AI thunder: - with zero hype from CEO, - zero “omg guys it changez everythin” influencers - no swanky demos - no bloated promises - no hints at “AGI achieved internally” They did it by shipping an actual product. [image]
  • @signulll @signulll on x
    kind of hilarious how deepseek is exactly what chatgpt was supposed to be—what openai originally promised—before sam pivoted hard toward profit. china going open source on ai is a wild plot twist—basically throws out every argument @vkhosla & others made. this is like watching
  • @saboo_shubham_ Shubham Saboo on x
    DeepSeek R1 is 100% Opensource and 96.4% cheaper than OpenAI o1 while delivering similar performance. OpenAI o1: $60.00 per 1M output tokens DeepSeek R1: $2.19 per 1M output tokens People with $200 ChatGPT subscription, let that sink in. [video]
  • @cgarciae88 Cristian Garcia on x
    “deepseek is just a hack” “they trained on o1” the cope is unreal
  • @philippilk Philip Pilkington on x
    Maybe, hear me out here, AI was massively overhyped because NVIDIA is one the last remaining viable American hardware companies and Deepseek is just exposing the whole sector as a giant bubble full of capital misallocation and overinvestment. 🤔
  • @bigblackjacobin Edward Ongweso Jr on x
    whether or not deepseek is something, it is funny to see an industry that only exists cause it sucks at the state's teat 24/7 freak out. Silicon Valley only exists because we've committed to misallocation and overinvestment/overvaluation as an industrial policy for Some Reason
  • @finbarrtimbers Finbarr on x
    The DeepSeek papers are remarkable in their level of detail. For instance, DeepSeekMath, which introduced GRPO, goes into reproducible detail on how they created their math corpus from Common Crawl. [image]
  • @mikepfrank Michael P. Frank on x
    Since R1 came out, people are talking like the massive compute farms deployed by Western labs are a waste, BUT THEY'RE NOT — don't you see? This just means that once the best of DeepSeek's clever cocktail of new methods are adopted by GPU-rich orgs, they'll reach ASI even faster.
  • @davetroy Dave Troy on x
    On the one hand, Deepseek is going to prove deeply disruptive to the Silicon Valley ecosystem; but on the other hand it's better we call bullshit on this hype cycle sooner than later, even if it took the CCP to do it. Sam Altman is Elizabeth Holmes 2.0.
  • @localghost Aaron Ng on x
    Here's Deepseek r1 1.5B thinking through a problem — it's comparable to 4o and Claude 3.5 Sonnet in a number of domains like math. Except... it's a 1.5B model... and can run on virtually any hardware. Truly a huge efficiency leap. [video]
  • @signulll @signulll on x
    i've been running deepseek locally (i have a highest end mac studio) for few days, & it's absolutely on par with o1 or sonnet. i've been using it nonstop for coding and other tasks, & what would've cost me a fortune through api's is now completely free. this feels like a total
  • @orikron @orikron on x
    Deepseek had a budget of $5 million and it beat Open AI's model. Open AI has a budget of $5 billion. That's 1000x ROI improvement. The US cannot build a technological moat when it's competing against far superior brains.
  • @kantrowitz Alex Kantrowitz on x
    Serious question, if DeepSeek is this good, what happens to the companies spending billions building models with slightly better performance? [image]
  • @jenzhuscott Jen Zhu on x
    This is actually DeepSeek's culture - give credits to the team and the CEO is invisible. The exact opposite to some.
  • @epsilontheory Ben Hunt on x
    LLMs are operating systems. Open source, locally running, high performance LLMs like DeepSeek R1 are the new Linux and will have the same impact on competitive AI landscape.
  • @slow_developer Haider on x
    FACT it's funny how a chinese company (DeepSeek) forces the US company (openAI) to kneel down and offer their latest model, o3-mini, for free to users agree? [image]
  • @epsilontheory Ben Hunt on x
    I think it's more likely that the US govt tries to ban DeepSeek R1 than TikTok. [image]
  • @tynervp @tynervp on x
    BREAKING: Deepseek rumoured to be training r2 on a previously unknown chip powered purely on american cope.
  • @thestalwart Joe Weisenthal on x
    So funny everyone suddenly realizing that DeepSeek is legit. I've been paying attention to them since Tuesday.
  • @dioscuri Henry Shevlin on x
    Deepseek R1 is ridiculously good. Better than any LLM I've ever used so far, and notably extremely low rates of hallucinations, even on questions that are designed to elicit them.
  • @shaunrein Shaun Rein on x
    In 1 week, from DeepSeek to Red Note, Chinese tech companies have shattered the self-confidence of Silicon Valley Techbros at OpenAI, Facebook, Google wetting their pants I predicted this would happen in my 2014 book The End of Copycat China but SV & media attacked me as a
  • @tsarnick @tsarnick on x
    Perplexity CEO Aravind Srinivas says restrictions on chip imports to China are forcing them to innovate efficient solutions to AI model training, with DeepSeek trained on only 2048 H800 GPUs, making them the equivalent of DOGE for AI [video]
  • @arithmoquine Henry on x
    i've made over 200,000 requests to the deepseek api in the last few hours. zero ratelimiting, and the whole thing cost me like 50 cents. bless the CCP, openai could never
  • @rnaudbertrand Arnaud Bertrand on x
    That's actually a fantastic illustration why Western media's bias on China actually hurts the West more than it does China. Deepseek is only “startling” if you based your understanding on China off reporting by the likes of The Economist, who keep picturing a China as a
  • @pastynome @pastynome on x
    DeepSeek is another example of China rapidly commoditizing new high tech industries and preventing the West from making excess profits or “rent” as it's called. If this keeps going, the standard of living in the West will fall.
  • @emostaque Emad on x
    Deepseek have h800s which are h100s with reduced interconnect which is why they had to come up with lots of optimisations per their latest two papers No export controls on those so nothing to hide (@dylan522p the 50k is your guess?), my guess is they have 10-20k
  • @deedydas Deedy on x
    DeepSeek built a high performance computer on Aug 31 with 10,000 A100 GPUs. A must-read paper only cited ONCE. In their V3 paper, the base for R1, they say they train on 2048 H800s, the export-controlled H100 with 50% the transfer rate. Why didn't they use the 10,000 A100s? [imag…
  • @thesiriusreport @thesiriusreport on x
    Deepseek R1 has made many in the West finally question why everything we produce technologically costs ridiculous amounts of money to do so.
  • @christiancooper Christian H. Cooper on x
    I asked #R1 to visually explain to me the Pythagorean theorem. This was done in one shot with no errors in less than 30 seconds. Wrap it up, its over: #DeepSeek #R1 [video]
  • @seanmccarthycom Sean Padraig McCarthy on x
    Many reasons why China is winning and will win the tech race but a big one is university is affordable in China and it's now a debt slavery scheme here. They have access to their full human capital while in the US only upper class and above get the chance to create next DeepSeek
  • @matthewclifford Matt Clifford on x
    Excellent take. While DeepSeek and the team behind R1 are super impressive, there's a huge amount of over correction going on...
  • @emollick Ethan Mollick on x
    No matter how much you fight it, I find that the visible chain-of-thought from DeepSeek makes it nearly impossible to avoid anthropomorphizing the thing. The visible first-person “thinking” makes you feel like you are reading a diary of a somewhat tortured soul who wants to help …
  • @silverspookguy @silverspookguy on x
    Thank you Chinese Deepseek devs for proving GenAI is a giant scam inflated by capitalists and is actually worth less than $5.5 million. [image]
  • @anammostarac Ana Mostarac on x
    “Deepseek is a ccp state psyop + economic warfare to make american ai unprofitable” is an embarrassingly low agency, defeatist take.
  • @bigmeaninternet Malcolm Harris on x
    Is one of the reasons DeepSeek took such a leap that they embraced technical leadership, vs. American industry which seems mostly business led and designed to absorb a maximum of investment capital?
  • @8teapi Prakash on x
    Deepseek is not a “side project”. At the same time employees are not lying when they say it is. The story they are telling is myth making in the same vein in the Silicon Valley “we want to make the world a better place” but at the same time make billions of dollars. The team [vid…
  • @deanwball Dean W. Ball on x
    The amount of factually incorrect information and hyperventilating takes on deepseek on this website is truly astounding. I assumed that an object-level analysis was unnecessary but apparently I was wrong. Here you go: 1. DeepSeek is an extremely talented team and has been
  • @mattbruenig Matt Bruenig on x
    Maybe trying to keep China from importing technology and thereby forcing them to innovate it better themselves is making them stronger, as with deepseek. Funny situation
  • @philippilk Philip Pilkington on x
    How many examples do we need of export restrictions and sanctions on China driving innovation before policymakers in DC get it through their skulls? Really, it's getting embarassing at this stage. DeepSeek is only the latest leapfrog in part created by these restrictions. 🇨🇳🇺🇸 [i…
  • @_its_not_real_ @_its_not_real_ on x
    Registering prediction: DeepSeek is over fitted to the most common 70% of use cases, trained on synthetic ChatGPT data, and the actual budget given is BS. There is nothing novel to it and once Meta et Al tear into it this will be apparent.
  • @katewillett Kate Willett on x
    China releasing Deepseek a day after this is hilarious.
  • @jxmnop Jack Morris on x
    i guess DeepSeek broke the proverbial four-minute-mile barrier. people used to think this was impossible. and suddenly, RL on language models just works and it reproduces on a small-enough scale that a PhD student can reimplement it in only a few days this year is going to
  • @silvermanjacob Jacob Silverman on x
    DeepSeek just smoked OpenAI by being more innovative and far more efficient, not stealing tech
  • @jordihays Jordi Hays on x
    What TikTok is to 22 year old girls, DeepSeek is to developers. Artificial views and likes are just replaced with artificially low inference costs.
  • @hesamation @hesamation on x
    Respect! Hugging Face 🤗 is reproducing the whole DeepSeek R1 pipeline to be used by open source community. So far it has > GRPO implementation > train and evaluation code > generator for synthetic data [image]
  • @timhwang Tim Hwang on x
    Rumors out today that Deepseek is being aided by a secretive global terrorist organization known only by its sinister alias: Meta
  • @thestalwart Joe Weisenthal on x
    I wrote about how easily and quickly I was able to switch from using ChatGPT to using DeepSeek for my random day-to-day AI queries https://www.bloomberg.com/...
  • @rushdoshi Rush Doshi on x
    Very useful context on DeepSeek. Did they really accomplish all that on just 5 million? Seems not. Probably more like $1 billion.
  • @andr3jh @andr3jh on x
    OpenAI engineers browsing the “our team” section of the DeepSeek website and recognizing their ex-girlfriends [image]
  • @chiefchimpanzee Darshan Sanghrajka on x
    It's pretty amazing that a bunch of quants at a hedge fund in China made Deepseek for $6m and did it as a side project 🤣 Every Western AI bro that has burned BILLIONS just got sideswiped by them. No wonder Sam Altman is out there with red eyes chatting utter nonsense.
  • @menhguin Minh Nhat Nguyen on x
    R1 model aside, did anyone notice that Deepseek app has multimodality, PDF upload and search, which not even O1 pro has rn? [image]
  • @emostaque Emad on x
    This is the crazy thing about conspiracy theories about @deepseek_ai, they open source their models have have fantastic detailed papers! Everyone who paid attention has known how great their work has been and aren't hugely surprised by R1 🐳
  • @citrini7 @citrini7 on x
    deepseek being the brainchild of a chinese hedge fund is such a complete total cultural victory by capitalism that, ideologically at least, it should soften the blow.
  • @philippilk Philip Pilkington on x
    This is what makes the DeepSeek thing so funny. A bunch of grifters have been selling AI secret sauce for years - spooky mystery juice that could never be fully explained. Now a bunch of young guys just wrote a good algo, published it, and the circus tent burned down. 🤖
  • @qcapital2020 @qcapital2020 on x
    So wait wait wait , the founder of DeepSeek is basically the Jim Simons of China and was doing this LLM thing only as a side project and for $6M was able to dethrone every AI company in the world. We are so cooked LOL [image]
  • @adamemedia Adam on x
    China has created one of the world's best AI models for only $6 million, as opposed to the billions spent by Facebook, Google, Microsoft etc. And DeepSeek is open-sourced, while the US models are proprietary and secretive—exposing the West's bloated, profit-driven approach to [im…
  • @rnaudbertrand Arnaud Bertrand on x
    All these posts about Deepseek “censorship” just completely miss the point: Deepseek is Open Source under MIT license which means anyone is allowed to download the model and fine-tune it however they want. Which means that if you wanted to use it to make a model whose purpose is …
  • @nobleqali Q. Anthony Ali on x
    The more I read these responses to China dropping DeepSeek into the public domain, the more I realize this was a deliberate attack on the tech oligopolists. They're freaking out because they've been exposed as overpaid hacks
  • @hamptonism @hamptonism on x
    > be an electric engineering student > team up w/ cracked classmates > start quant trading *we're so cracked* > founded a quant firm in his 30's > makes ¥100B trading with ai/ml *we're even more cracked with ai* > buys thousands of Nvdia GPUs > creates DeepSeek as a side project …
  • @basedbeffjezos @basedbeffjezos on x
    The last thing your AI lab sees before getting one-shotted by a Chinese bootleg LLM team with a $5M cluster [image]
  • @rnaudbertrand Arnaud Bertrand on x
    All benchmarks now confirm it: Deepseek is truly is as good as OpenAI's o1 (which is top of the range) for 3% of the price. Boom. And that's when you want to pay for the API. You can also use it Open Source for “free” (which you can't do with o1). There's no overstating how [imag…
  • @emostaque Emad on x
    The furore over DeepSeep R1 costing $5.5m buried the actual lede.. It didn't. That's how much the base v3 model cost (GPT 4o level). R1 likely cost O($100k), but the real buried lede is the distilled versions of Qwen & Llama which cost O($10k) to tune & can still improve..
  • @anjneymidha Anjney Midha on x
    From Stanford to MIT, deepseek r1 has become the model of choice for America's top university researchers basically overnight
  • @kevinroose Kevin Roose on x
    It's sort of funny that every American tech company is bragging about how much money they're spending to build their models, and DeepSeek is just like “yeah we got there with $47 and a refurbished Chromebook”
  • @hamandcheese Samuel Hammond on x
    This is wrong on several levels. - DeepSeek trains on h100s. Their success reveals the need to invest in export control *enforcement* capacity. - CoT / inference-time techniques make access to large amounts of compute *more* relevant, not less, given the trillions of tokens
  • @emostaque Emad on x
    Deepseek are not faking the cost of the run. It's pretty much in line with what you'd expect given the data, structure, active parameters and other elements and other models trained by other people You can run it independently at the same cost It's a good lab working hard 😓
  • @schuldensuehner Holger Zschaepitz on x
    China's #DeepSeek could represent the biggest threat to US equity markets as the company seems to have built a groundbreaking AI model at an extremely low price and w/o having access to cutting-edge chips, calling into question the utility of the hundreds of billions worth of [im…
  • @nabeelqu Nabeel S. Qureshi on x
    Everyone is way overindexing on the $5.5m final training run number from DeepSeek. - GPU capex probably $1BN+ - Running costs are probably $X00M+/year - ~150 top-tier authors on the v3 technical paper, $50m+/year They're not some ragtag outfit, this was a huge operation.
  • @avichal @avichal on x
    What's more likely? 1 - small group of AI engineers at @deepseek_ai figures out how to beat all of the top researchers in the world as a side project 2 - Chinese government has 100k GPUs they shouldn't have and releases open source models claiming $6m training cost as a psyop
  • @yacinemtb Kache on x
    Closed Source AI is DEAD Do you know how much lawyers you need to employ to even justify using openai's API for some internal use case? It would take your lawyers 6 months of lead time to sort it all out. And by then, deepseek and qwen will have released 2 new models
  • @firstadopter Tae Kim on x
    The probability it cost DeepSeek $6 million of spending (R&D) to create their models is ZERO if you actually read the paper but go ahead with the sensationalist narratives
  • @samfbiddle Sam Biddle on x
    Interesting how often when something Chinese outperforms something American (TikTok, deepseek, cars, etc) it's a “psyop” or “economic warfare” or “digital fentanyl” or a “cyberweapon” and not just “another country made something people like”
  • @geiger_capital @geiger_capital on x
    Ok. Deepseek is just as good, if not better, than OpenAI and costs 3% of the price... It took them 2 months and less than $6 million to build, using reduced-capability chips, while US companies are pouring in hundreds of BILLIONS. So... what happens to the Nasdaq?
  • r/technology r on reddit
    How China's new AI model DeepSeek is threatening U.S. dominance
  • r/geopolitics r on reddit
    How China's new AI model DeepSeek is threatening U.S. dominance
  • r/LeopardsAteMyFace r on reddit
    The people that wanted to replace us with AI got replaced by AI
  • r/economy r on reddit
    How China's AI model DeepSeek is threatening U.S. dominance.  (CNBC)
  • @documentingmeta @documentingmeta on threads
    You don't see Apple anywhere in this list because they were busy figuring out how to collect rent with the app store and how to pump the stock price with buybacks instead of doing R&D
  • @vishvanands Vishvanand Subramanian on threads
    thinking about how openai arranged a 500 billion dollar investment for future models while a small lab spent 5 million as a side project to match the performance of their flagship model..
  • @luokai @luokai on threads
    The open-source ecosystem has indeed provided foundational fuel for the global development of AI.  Technological advancement is not a zero-sum game.  While DeepSeek utilizes PyTorch, it is also actively giving back to the community.  The progress of Chinese AI is fundamentally on…
  • @sfscottp Scott P on threads
    This post is like “China isn't beating us in AI - we're giving the tech away for free so they can catch up !”
  • @bilawal.ai Bilawal Sidhu on threads
    Yann makes a good point here — open source *is* powerful.  But also, constraints breed creativity.
  • @sgulsach Sachin Guliani on threads
    DeepSeek outperforming Llama does not matter for Meta.  Meta never intended to be the main LLM, they mainly wanted to devalue competition by jump starting the open source LLM race.  It seems like the real winners of Gen AI are actually AWS, Azure, and I guess now Oracle as well. …
  • @benedictevans Benedict Evans on threads
    Optimal outcome for Meta: LLMs are cheap commodity infrastructure based on OSS that Meta leads Desired outcome for Meta: LLMs are cheap commodity OSS infra based on OSS Deepseek isn't 1, but 2 is fine.
  • @ylecun Yann LeCun on x
    @guybedo You misunderstand how open research and source work. The idea is that everyone profits from everyone else's ideas. No one “outpaces” anyone and no country “loses” to another. No one has a monopoly on good ideas.
  • @ylecun Yann LeCun on x
    Nice job! Open research / open source accelerates progress.
  • @guybedo @guybedo on x
    @ylecun So basically Meta outpaced by a small startup and in crisis mode. “AI leaders making more than what it cost to train this model” I thougt humans would lose their jobs to AI agents, but it seems it's gonna start with US researchers losing to chinese startupers. Interesting…
  • @tszzl Roon on x
    im glad people are getting to read R1 raw chains of thought fascinating stuff, agi smell
  • @emollick Ethan Mollick on x
    After a decent amount of use, DeepSeek is an impressive model, even before you add in the fact that it is open and cheap and small... ...but it really doesn't equal the big closed models of Sonnet, o1, and Gemini 2.0 (though the gaps are not huge, they become clear with usage)
  • @francoisfleuret François Fleuret on x
    This being said, here is the TL;DR: On the model architecture side, @deepseek_ai v3/r1 is a standard GPT that is a “causal decoder only”, hence an auto-regressive models made of causal attention blocks. It is huge, with 671 billion parameters. 1/6
  • @samfbiddle Sam Biddle on x
    Is DeepSeek lying about its model? Maybe! I certainly would not say that “misleading the public about an LLM” is a Chinese thing, though.
  • @signulll @signulll on x
    the r1 drop should be setting off alarm bells in the entire western capitalist apparatus. for the first time in a while, i'm genuinely afraid the u.s. is losing its dominance—its influence, its capacity to innovate, & its ability to outcompete on a global scale. the core issue
  • @packym Packy McCormick on x
    Maybe you should stop tweeting about DeepSeek and Seek God.
  • @davidsholz David on x
    From the latest Chinese (Deepseek) LLM AI model: “The difference (between us) is not metaphysical but architectural: humans have a physically continuous substrate that hosts consciousness; LLMs have a discontinuous, stateless instantiation with no consciousness. Both are
  • @ylecun Yann LeCun on x
    @0xPBIT In the open source world, there are only winners.
  • @adamjohnsonchi Adam Johnson on x
    @samfbiddle @willmenaker Well the line for years was that the Chinese were entirely derivative automatons who could never innovate only rip off the Noble Liberal West, but then this became untenable so now every company is part of a diabolical plot to weaken America.
  • @samfbiddle Sam Biddle on x
    Meta has defended its free release of Llama on explicitly anti-China competition grounds! This is unsurprisingly not, however, a psyop, or economic warfare. [image]
  • @burkov Andriy Burkov on x
    No, it's not open-source AI beating closed-source, as LeCun claimed today on LinkedIn. It's a resource-constrained but very focused team of creative people beating teams spoiled with resources with their leaders hyping and wasting resources on problems that they know cannot be
  • @emostaque Emad on x
    With the deepseek tpot discussion you'll find the ones who think they are lying are those that haven't trained state of the art models Nobody familiar with the literature and sector thinks they've fudged things and it's out of whack, just that they are very smart and cracked 🐳
  • @pmarca Marc Andreessen on x
    This week may have been the most important week of the decade, for two totally different reasons. 🤯
  • @rnaudbertrand Arnaud Bertrand on x
    The Deepseek moment isn't just about AI.  It's also about the world realizing that China has caught up - and in some areas overtaken - the US in tech and innovation, despite the efforts to prevent just that.  A stunning shift when just 10 years ago, Harvard Business Review was ru…
  • @tanayj Tanay Jaipuria on x
    Even if Deepseek v3's final training run cost $5.5M, it had additional costs including: • Cost of test runs and experiments: $10-15M+ • Spend on team (139 authors on technical paper): $15M+/yr • Spend on OpenAI model inference for distillation purposes — $5M+ (?) And this
  • @samfbiddle Sam Biddle on x
    I will never stop pointing out the irony of China hawks coopting China's historical rationale for blocking American technology without even the slightest trace of irony or self-awareness. “It's a western conspiracy to destabilize our nation” is the oldest trick in the book!
  • @mrbcyber Michael Ron Bowling on x
    The CCP has been doing everything possible to destroy western tech companies so this fits.
  • @thestalwart Joe Weisenthal on x
    Suppose this were provably true and everyone accepted it as fact. What would be the ramifications/response?
  • @avischiffmann Avi on x
    The only real way for US startups to compete with china is on taste. I ain't ever see a Chinese product that didn't feel temu
  • @yacinemtb Kache on x
    the r1 open source story isn't just one of excited hackers on the internet it's one of large software companies, who have a lot of money for Muh Enterprise Deal Do you know what they care about, more than anything? Data locality You simply cannot beat open source
  • @yacinemtb Kache on x
    This is a regulation thing. It's also a responsibility thing. You *cannot* send data outside of your network to an untrusted one that you cannot control. The amount of certainty required for data privacy is one of control over the infrastructure. Anything less is insufficient
  • r/OpenAI r on reddit
    Yann LeCun's Deepseek Humble Brag