/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

Kimi K3 hype shouldn't alarm the US about “losing the AI race” to China; K3 is a good model but not frontier-level and likely lacks dangerous cyber capabilities

Transformer Weekly: NY data center moratorium, NDAA export controls and Amodei's $1m to safety super PAC

Transformer

Context & Ripple Effects

Coverage of Kimi K3 moved quickly from reports that Moonshot was preparing a very large release to benchmark results placing it near leading coding systems. This article adds a limiting interpretation: strong task performance is not, by itself, evidence of frontier status or dangerous cyber capability.

That distinction matters amid linked debates over export controls, data-center buildout, and AI safety politics: the policy stakes depend on capabilities and access, not on a single leaderboard rank.

First-order effects

  • Kimi K3's reported benchmark showing strengthens Moonshot's credibility with developers and buyers evaluating coding models, while the article argues against treating it as proof that China has overtaken U.S. frontier labs.
  • Policymakers and safety advocates have a clearer basis to separate competitive concern from claims about immediate high-end cyber risk.

Second-order effects

  • Rival labs face more pressure to demonstrate performance across varied evaluations, especially as coding benchmarks can rapidly shape public perceptions of model leadership.
  • Export-control and AI-safety debates may shift toward more capability-specific thresholds; broad reactions to a prominent benchmark result become harder to justify on their own.

Third-order effects

  • If capable but non-frontier models continue to narrow visible benchmark gaps, leaderboard leadership will become a weaker proxy for either national AI leadership or catastrophic-risk exposure.
  • The durable contest is likely to center on who can sustain frontier capability, compute access, deployment infrastructure, and credible safety assessment—not simply who wins a coding evaluation.

The trend: AI geopolitics is moving from headline model rankings toward more granular judgments about capability, infrastructure, access, and risk.

Discussion

  • @emollick Ethan Mollick on x
    Kimi is, as I have been saying, a very good model. But it is not a DeepSeek r1 moment, in that it is roughly where I would expect on the curve rather than an unexpected leap. It will be treated as a DeepSeek moment for a variety of reasons especially as more people hear about it.
  • @emollick Ethan Mollick on x
    A lot of swift conclusions are being drawn about Kimi K3 based on fairly saturated benchmarks and ELOs, rather than actually testing it on very hard problems. The AI frontier has already moved so far that a good model that is a still months behind looks like the future to many.
  • @shakeelhashim Shakeel on x
    Kimi K3 threatens to once again send Washington into a panic. But look a little closer, and the hysteria may be unwarranted. On Thursday, the AI industry got its second “DeepSeek moment” — this time courtesy of Moonshot, whose new Kimi K3 model has “erased America's AI lead,”
  • @zephyr_z9 @zephyr_z9 on x
    After the K3 drop, Huawei has now released a real 950 SuperPoD 1 EFLOPS fp8 & 2 EFLOPS fp4 256TB of unified memory [image]
  • NewsMax.com Charlie McCarthy on x
    China's New AI Threatens US Tech Lead
  • @deanwball Dean W. Ball on x
    I wonder if the California attorney general or the European Union ai office will seek to make moonshot comply with their respective frontier ai regulations
  • @shakeelhashim Shakeel on x
    This is an inaccurate and irresponsible headline from Axios. [image]
  • @levie Aaron Levie on x
    This post is key. The cheaper AI gets, the more opportunity there is for the entire ecosystem - especially including end-customers - to benefit. Everything is bottlenecked by being able to successfully and cost effectively deploy AI in real workloads. Any time we can lower the
  • @pstasiatech Paul Triolo on x
    ...the hypothesis that they have much more advanced model checkpoints internally that are already being used for RSI. In the latter scenario, reaching RSI even a few months ahead of other labs might be enough to cement a permanent lead.
  • @notthreadguy @notthreadguy on x
    genuinely one of the best bull posts I've ever read > China open source ai is better than US? Long all capex beneficiaries > China open source ai is NOT better than US? Long all capex beneficiaries
  • @shanumathew93 Shanu Mathew on x
    Great insights. Gavin nails it. If an oligopoloy of labs sustain 90% inference margins, they capture most of the economics and eventually vertically integrate the stack & squeeze chips, power, data centers, cloud and software. More competition at the model layer —> lower model
  • @bare_birk Birk on x
    Interesting post from @GavinSBaker! Kimi K3 is bad for Anthropic and OpenAI but good for all other companies. Margins will go from the frontier labs, to all other companies in the sector. Infrastructure will still be very important in both scenarios (Opensource vs closed source)
  • @chamath Chamath Palihapitiya on x
    Gavin is right. This is positive for everyone except the closed frontier labs.
  • @samanthaladuc Samantha LaDuc on x
    A bullish Kimi K3 argument: “Anything that lowers margins and increases competition at the model layer is good for every other AI layer: power, semiconductors, hyperscalers, neoclouds and yes even software.”
  • @dzambhalahodl Steven Lubka on x
    Kimi3 is a good model, but it is not a “cheap” model nor is it a “small” model. Kimi3 showed that China can train a decent competing model less than 6 months behind the US. It did not show that they can compete with Frontier at a similar efficiency advantage to Deepseek etc.
  • @liangsays Brent Liang on x
    one of the best takes on this site re kimi. we are chronicling a mini sputnik moment right now and living through history
  • @limitingthe @limitingthe on x
    “Anything that lowers margins and increases competition at the model layer is good for every other AI layer: power, semiconductors, hyperscalers, neoclouds and yes even software.”
  • @omarsar0 Elvis on x
    That missing token efficiency is coming. Can't say more now, but efficient frontier long context reasoning/understanding and other TTC breakthroughs are on the horizon. Architectural improvements are often ignored in these benchmarks, but I expect major shifts by EOY.
  • @arena @arena on x
    Big news: Kimi-K3 by @Kimi_Moonshot is now #1 in the Frontend Code Arena with 1679 pts, surpassing Claude Fable 5. This is a 17-place jump from Kimi-k2.6 (#18 -> #1). In Frontend, Kimi-K3 ranked #1 in 6 of 7 domains: Brand & Marketing, Reference-Based Design, Data & Analytics, [i…
  • @artificialanlys @artificialanlys on x
    Kimi K3 scores 57 on the Artificial Analysis Intelligence Index. Its intelligence is comparable to Opus 4.8 and GPT-5.5 but remains behind Fable 5 and GPT-5.6 Sol. Moonshot AI has expressed plans to release the 2.8T parameter model's weights, which would make it the leading open …
  • @davidsacks David Sacks on x
    This is concerning. For the first time, a Chinese model Kimi K3 has taken #1 on the Frontend Code Arena and is scoring at or near the frontier on other benchmarks. Meanwhile America is tying itself in knots: politicians and bureaucrats are banning new data centers, piling on
  • @quxiaoyin Xiaoyin Qu on x
    What Kimi K3 means for USA AI: 1. FYI Kimi K3 is open weight and will be released on July 27, 2026. 2. When the best open weight model exceeds the best closed-source model, how does @AnthropicAI justify its Fable pricing? Why would anyone pay for that? Not to mention the Mythos
  • @ryangreenblatt Ryan Greenblatt on x
    Kimi K3 was significantly but not massively above my expectations. I'd tentatively guess it's similar in overall usefulness/usability to Opus 4.8 and in overall capability somewhat above Opus 4.8 (while also being somewhat more benchmaxxed). As a pretrain, it's probably somewhere
  • @arena @arena on x
    Kimi-K3 just topped the Frontend Code Arena with a 76% pairwise win rate. When its output was compared head-to-head against other models on the same task, it was picked as the better output 76% of the time on average. For reference: Claude Fable 5 (63%), GPT-5.6 Sol (58%). 50% [i…
  • @mweinbach Max Weinbach on x
    It finally finished, here's the final output. Used 60% of my monthly Kimi usage on it https://macos27.kimi.page/
  • @pimdewitte Pim de Witte on x
    kimi is truly a revolutionary model [image]
  • @aftfuture @aftfuture on x
    .@DavidSacks is right. America is not going to regulate, restrict, and delay its way to victory in the AI race. Blocking data centers, piling on state mandates, and requiring federal preapproval for frontier models would not make us safer. It would make China stronger.
  • @benbajarin Ben Bajarin on x
    Exactly right. You can argue it's bad for the frontier model labs but all lower token costs does is drastically increase demand for compute. It's the hyperscalers who benefit from this more than anyone and even if THEY were they only ones spending, there will still be more
  • @mweinbach Max Weinbach on x
    I asked a Kimi K3 Max agent swarm to recreate macOS 27 with real Liquid Glass and native apps in web browser and it's been going for 3 hours [image]
  • @nicky_sap @nicky_sap on x
    Everyone raving about Kimi K3 so I had to give it a shot. I signed up for free and threw it a pretty nebulous prompt: 'look up Nick Saponaro and create a groundbreaking website using animation, cursor reaction, masterful design, and unique user experience. not just a parallax [vi…
  • @obrien Chris O'Brien on x
    In retrospect, was it perhaps a mistake to elect a guy whose primary concern was getting eaten by sharks while on an electric boat to lead the country through this period?
  • @miles_brundage Miles Brundage on x
    K3 analysis wen Also, reminder to Americans - we could have this kind of state capacity at home. Let's properly fund and unmuzzle CAISI! https://x.com/...
  • @laschuk @laschuk on x
    I put Kimi K3 up against Fable 5 in my benchmark. The test: clone Apple's homepage. Single prompt, one shot, zero help. Kimi K3: $0.44 (at left) Fable 5: $0.94 (at right) Same test. Same prompt. Half the price. Fable runs $10/$50 per MTok, the most expensive tokens on the [video]
  • @bhavani_00007 @bhavani_00007 on x
    I tested Kimi K3 vs Claude Opus 4.8 Same prompt, an armory bay with lighting, props, and detail. Top is Kimi K3, bottom is Opus 4.8. It's not even close. Kimi K3 built a full scene with textures, proper lighting, ammo crates, weapon racks, working detail everywhere. Opus 4.8 [vid…
  • @aisecurityinst @aisecurityinst on x
    Our first public analysis of the open/closed weight gap in frontier cyber capabilities finds it is 4-7 months with GLM-5.2 and DeepSeek V4-Pro, narrowing from 6-10 months through most of 2025. Advanced capabilities are reaching less safeguarded open models faster than before. 🧵 […
  • @jiahanjimliu Jim Liu on x
    Neoclouds: The Kimi K3 Scare Kimi K3 caused a large scare in the AI trade as this Chinese open source model matched frontier models on benchmarks. Let me unpack what's actually going on. Chinese Labs have much less GPUs than American Labs and yet are able to train “just as
  • @nic_carter Nic Carter on x
    From an ecological perspective, western frontier models are a common pool resource that Chinese open weight models exploit. They are finite, because if you exploit them too hard, they lose the incentive and resources to train new models, and everyone loses. (Chinese models aren't
  • @jukan05 Jukan on x
    Anyone loudly touting how much cheaper Kimi K3 is than Western LLMs is deliberately ignoring how much more expensive it has become compared with previous Chinese LLMs.
  • @alexfinn Alex Finn on x
    I was wrong. I said we were a year away from Fable 5 on our desk. That day is today An open model BETTER than Fable 5 in some benchmarks just dropped Better than ChatGPT 5.6 on FrontierSWE. Better than Fable 5 on Automation Bench This fundamentally changes the AI race forever
  • @semianalysis_ @semianalysis_ on x
    Kimi K3 2.8T is so large that it will not fit on a single NVIDIA DGX B200, even at FP4. A GB300 NVL72, B300, or MI355X system is required, as each GPU has 288 GB of memory. One optimization that could make Kimi K3 fit on B200 is to gang multiple nodes together and use a [image]
  • @signulll @signulll on x
    a while ago there was a device called the palm pre, pretty revolutionary at the time. it ran something entirely new called webOS, which was built using html5. it had multi tasking, card based navigation, was very cool & slick. palm needed hundreds of engineers, years of
  • @arena @arena on x
    In the Text Arena, Kimi-K3 by @Kimi_Moonshot landed #9, with 1486 pts. This is another significant improvement from Kimi-k2.6 (#38 -> #9). - Top 10 in Creative Writing, Coding and Instruction Following - #1 in three occupations: Physical & Social Science, Legal & Government, [ima…
  • @deredleritt3r Prinz on x
    The most interesting question about Kimi K3 is whether it poses cyber risk. Kimi K3 benchmarks do not include a CyberGym score. Waiting for @AISecurityInst to bench this model.
  • @yzhang_cs Yu Zhang on x
    K3 has now crossed the 1M context-length barrier, and DeepSeek's sparse attn has done the same. But what architecture will take us to 5M, 10M, or even longer? I'd always argue that fixed-state linear attn, especially GDN/KDA, is highly competitive here. Hybrid designs are
  • @yzhang_cs Yu Zhang on x
    funny that K3 is great at making 3Blue1Brown videos, and the Quantile Balancing example in the blog took just a few shots to produce. https://kimi.com/...
  • @arena @arena on x
    Kimi K3 from @Kimi_Moonshot has moved the Pareto Frontier for Code Arena: Frontend. [image]
  • Georg Zoeller Georg Zoeller on linkedin
    The reason this all ends up with war one day is because to Americans everything everyone in the world does is about them and if they are not winning, it's always an attack. …
  • Dave Schroeder, PhD Dave Schroeder, PhD on linkedin
    This is intentional, by design, and strategically intended in every dimension to elicit exactly this panicked reaction, to destabilize US AI development …
  • @metacurity.com Cynthia Brumfield on bluesky
    China's Moonshot just issued its new Kimi K3 model, which the company said rivals the strongest offerings from OpenAI and Anthropic PBC, and this is freaking everyone out, even though it was utterly predictable.  —  www.bloomberg.com/news/article...
  • @carlquintanilla Carl Quintanilla on bluesky
    “.. People are worried that if US companies start using Chinese models more and Anthropic less, then Anthropic will invest less.  That means those US firms will lower the capex and in the end chip demand will be affected.”  —  @bloomberg.com  —  www.bloomberg.com/news/article... …
  • @carnage4life Dare Obasanjo on bluesky
    More coverage of Kimi K3's achievements below
  • r/technology r on reddit
    China's Moonshot unveils world's largest open AI model, closing in on US rivals
  • r/neoliberal r on reddit
    China's open-weight Kimi model stuns AI world with frontier-level results
  • r/StockMarket r on reddit
    China's open-weight Kimi model stuns AI world with frontier-level results
  • @deanwball Dean W. Ball on x
    Some observations on Kimi: 1. It's a very good model! I don't think its performance can be explained away by distillation or anything like that. In agentic coding sessions, it seems pretty much on par with the best public models of Q1 2026. In my fairly limited use, it also
  • @teortaxestex @teortaxestex on x
    Kimi K3 and GLM 5.2 are remarkable in that these labs *underreport* their model's performance on the most valuable evals in official announcements. In DeepSWE, K3 is not 2.5% behind Fable. More like 1%. And cheaper. We're a long way from “benchmaxxing” era. Update accordingly. [i…
  • @counternotions Kontra on x
    Three decades ago China's largest export category was clothing. Time flies.
  • @negligible_cap @negligible_cap on x
    The fact that $BABA owns 36% of Moonshot (is their biggest single investor) but dropped 4% overnight seems like evidence that there's a lot more than just Kimi to this tech selloff News on Moonshot's most recent valuation was in June, when they were looking to raise at a $30B
  • @willccbb Will Brown on x
    oh no what if the big labs can't easily recoup all of their capex on massive data centers and some of the buildout capacity gets offloaded to other providers who deliver it to the long tail of enterprises with a software stack for serving and continually improving open models
  • @zephyr_z9 @zephyr_z9 on x
    Damn Kimi recreated Windows XP
  • @zck Zak Kukoff on x
    A few thoughts about what might happen as open models continue to achieve near-frontier performance: 1/ Closed-lab revenue is basically a function of distance between frontier outperformance and open model commoditization - call it six months. So the AI race accelerates
  • @xeophon Florian Brand on x
    1/5 the cost, less tokens and the same performance as Fable [image]
  • @rsalakhu Russ Salakhutdinov on x
    Congratulations to Zhilin Yang, founder and CEO of @Kimi_Moonshot, on the latest Kimi release. What a huge win for the open-source community! It feels like just yesterday Zhilin was graduating from my lab at CMU, jointly co-advised with William Cohen. Not only did he complete [im…
  • @agupta Ankit Gupta on x
    FYI that America's moronic visa policy is almost certainly a factor in why Kimi/Moonshot is a Chinese startup and not an American one. The fact that we don't staple a green card to every AI PhD completed in America is stupid. would be more logical to seize their passports and
  • @ctjlewis Lewis on x
    Count the number of competitors YC has put up to Kimi or DeepSeek and then let me know if you find hawkish immigration policy across three administrations to be the real problem here.
  • @theo @theo on x
    Kimi k3 is an incredible model. It is not an incredible value. In most tasks, it comes out to roughly the same cost as GPT-5.6 Sol. K3 is half the price of 5.6 Sol per token. GPT-5.6 uses half as many tokens. Price evens out. GPT-5.6 is 2x faster TPS, so it gets work done ~4x [im…
  • @enzo_gte Enzo on x
    I've been using Kimi K3 for ~16 hours now. The model is clearly good at a lot of different things (especially frontend), but non obvious reason why people are enjoying it so much is that it clearly does not follow the same rules in terms of safeguards and copyright. Kimi will [im…
  • @semianalysis_ @semianalysis_ on x
    CHINA'S KIMI K3 HAS SURPASSED ALL AMERICAN MODELS IN FRONT-END CODING WHILE BEING SMALLER THAN MOST CLOSED-SOURCE FRONTIER MODELS. Great work by the @Kimi_Moonshot team. [image]
  • @crystalsssup Crystal on x
    Kimi K3 makes me feel like one prompt is enough to build a game. You can build a game console, plug in a cartridge, and play a game on it. I just played one of my fav game “Ace Attorney” on my computer! [video]
  • @rajaxg Raja Koduri on x
    (K)impressive! As an early proof of concept, Kimi K3 designed a chip to serve a nano model built on its own architecture. In a single 48-hour autonomous run, K3 built, optimized, and verified the chip using open-source EDA tools on the Nangate 45nm library. Within 4 mm², the
  • @datacurve @datacurve on x
    Kimi K3 debuts at #3 on DeepSWE. It's the first open-weights model that delivers frontier-level performance, achieving results similar to Claude Fable and GPT-5.6 Sol. [video]