Analysis: recent open weight models lag frontier closed models' cyber capabilities by 4 to 7 months, a narrower gap than the 6 to 10 months through most of 2025
AISI has tracked the cyber capabilities of frontier AI models since 2023. On our evaluations, the most capable models …
AI Security Institute
Context & Ripple Effects
AISI’s cyber-range coverage has already shown that capability is uneven even among frontier systems: Mythos Preview completed both ranges while GPT-5.5 completed one. Separately, NIST assessed DeepSeek V4 Pro as roughly eight months behind leading US models, placing model-release lag at the center of comparative capability analysis.
The new measurement indicates that the open-weight versus closed-model gap is compressing from the range AISI observed through most of 2025. That matters because cyber capability is no longer confined as long to providers that control access to their models.
First-order effects
- Developers and users of recent open-weight models gain access to cyber-relevant capability closer to the closed-model frontier, reducing the practical lead held by closed-model providers.
- Defenders and AI safety teams must treat open-weight releases as a more current input to cyber-risk assessments, rather than relying on older assumptions about their capability lag.
Second-order effects
- Closed-model vendors face more pressure to differentiate through agent performance, reliability, tooling, and controlled deployment—not simply a capability lead—when open models catch up faster.
- Access governance becomes harder to base on a closed-versus-open distinction alone: a shorter lag means the same evaluation results increasingly matter across both distribution models.
Third-order effects
- If the compression persists, frontier cyber capability may diffuse through the ecosystem more quickly after it appears in closed models, narrowing the time available for provider-led controls to be the primary safeguard.
- The durable industry split may shift from who has the strongest base model to who can evaluate, govern, and deploy cyber-capable systems responsibly; the limited set of reported evaluations leaves the pace and breadth of that shift uncertain.
The trend: This is a data point in the faster diffusion of frontier AI capabilities from controlled services into more widely deployable model weights, increasing the importance of model-access governance and evaluation capacity.
Related: Frontier-model access governance · Frontier-model concentration risk · AISI · Mythos Preview is the first AI model to complete both of AISI's cyber · An evaluation by NIST's CAISI says DeepSeek V4 Pro lags behind leading
Related Coverage
- Meet the Companies Shelling Out for Top AI Models Wall Street Journal · Belle Lin
- Chinese AI models narrow cyber gap with US rivals Financial Times
- How Far Behind the Frontier are Leading Open Weight Models on Cyber? Lobsters
- Why China's Kimi K3 has Silicon Valley — and Washington — quaking Business Insider · Thibault Spirlet
- China's Moonshot AI Unveils Kimi Model, Threatening America's Lead New York Times
- Chinese AI has leveled up, and brought renewed focus on the open weight model shift CNBC
- Meet Kimi K3, the newest Chinese AI model haunting Silicon Valley Dow Jones Newswires
- The catch behind Kimi K3's benchmark leap The Deep View · Nat Rubio-Licht
- ✨ Down with American data centers! Up with Chinese AI! Faster, Please! · James Pethokoukis
- There's One Way To Win The AI Race, And The Big Labs Are Lobbying Against It Forbes · Christian Catalini
- Maybe you don't need the brand name AI Tech Brew · Whizy Kim
- Kimi K3 Just Triggered DeepSeek Flashbacks for the Stock Market Decrypt · Jose Antonio Lanz
- Moonshot AI's Latest Model Dings Tech Stocks. We've Seen This Story Before. Barron's Online · Nate Wolf
- Kimi K3 threatens AI business models Semafor · Reed Albergotti
- Moonshot reveals new AI model, and it's a big surprise — here's why Kimi K3 is a threat to the likes of OpenAI TechRadar · Darren Allan
- China's Moonshot AI Unveils Kimi K3, a Massive Parameter Model Challenging U.S. Dominance Techstrong.ai · Jon Swartz
- Report: China's ‘Kimi K3’ AI Model Matches Performance of Top American Systems Breitbart · Lucas Nolan
- China Just Unveiled An AI Model That Rivals America's Best. Silicon Valley Is No Longer Clearly Leading The Race. International Business Times · Merin Rebecca Thomas
- Chinese AI firm Moonshot unveils powerful model with capabilities close to Anthropic, OpenAI New York Post · Thomas Barrabi
- China's Moonshot launches Kimi K3 open AI model Verdict · Anwesha Pattanaik
- Moonshot AI unveils Kimi K3, the world's largest open-weight AI model: What to know The Indian Express
- AI Model Prices Are Falling At The Worst Moment For The U.S. Frontier Labs Big Technology · Alex Kantrowitz
- AI & Tech Brief: China's Moonshot AI Washington Post · Benjamin Guggenheim
- David Sacks, Bill Ackman Sound the Alarm on China's Kimi K3 as Nvidia, Micron Slide Benzinga · Daragh Thomas
- Chinese AI Startup Moonshot Unveils Kimi K3 Model—Will It Challenge OpenAI And Anthropic? Forbes · Ty Roush
- What smart people are saying about China's hot new Kimi K3 AI model Business Insider
- A New Blow to Tracking Gun Sales New York Times
- China's Moonshot Unveils AI Model, Fueling Tech Rout Bloomberg
- China's Moonshot AI launched its largest model ever — and rival AI stocks plunged Quartz · Cris Tolomia
- Moonshot AI's New Kimi K3 Challenges U.S. Frontier Models The Information · Juro Osawa
- China's Moonshot unveils world's largest open AI model, closing in on US rivals Reuters · Laurie Chen
- Kimi K3 Just Triggered DeepSeek Flashbacks for the Stock Market Yahoo Finance · Jose Antonio Lanz
- David Sacks challenges US AI policy after China's Kimi K3 tops coding test crypto.news · Lawrence Mondal
- China's Kimi K3 Is Out—And Beats Claude Fable and GPT 5.6 Sol on Key Benchmarks Decrypt · Jose Antonio Lanz
- New Chinese AI chatbot Kimi K3 rivals US leaders in the field Washington Examiner · David Zimmermann
- Kimi K3 Intelligence, Performance & Price Analysis Artificial Analysis
- China's Moonshot AI claims Kimi K3 can rival OpenAI and Anthropic BBC · Francisco Velasquez
- Moonshot's Kimi K3 Won't Fit on a Single Nvidia DGX B200 Implicator.ai · Marcus Schuler
- Kimi K3 tops Arena's coding leaderboard — and it's open-weight The New Stack · Amanda Caswell
- Former White House Crypto Czar David Sacks Warns US Could Lose AI Race CoinGape
- Kimi K3's performance bolsters Sacks' case against AI regulation Axios · Madison Mills
- China's Moonshot AI gains on Anthropic, OpenAI with new low-cost model Nikkei Asia · Itsuro Fujino
- Moonshot unveils Kimi K3, largest open-weight AI model yet Silicon Republic · Suhasini Srinivasaragavan
- Kimi (Moonshot AI)'s Post LinkedIn
- Bitcoin faces fresh headwinds as China's Kimi beats Claude, GPT in coding benchmark CoinDesk · Shaurya Malwa
- Kimi K3 Built A Chip In Just 48 Hours, Which Pushes Over 8700 Tokens/s, As China's Moonshot Delivers A 2.8 Trillion Parameter Frontier AI Model Wccftech · Hassan Mujtaba
- China's 2.8-trillion-parameter Kimi K3 beats Claude Fable 5 in Frontend Code Arena benchmark— Moonshot AI delivers largest open-weight AI model ever, as China works around U.S. compute limits Tom's Hardware · Luke James
- 😼 Kimi K3 goes open The Neuron · Grant Harvey
- Moonshot AI Unveils 2.8T-Parameter Kimi K3 AI Model WinBuzzer · Markus Kasanmascheff
- China's Kimi K3 AI Model May Trigger a DeepSeek Moment. Expect a Lot of FUD. Key Context · Tae Kim
- Moonshot's Kimi K3 closes the frontier gap The Rundown AI
- Sacks: China's latest frontier AI advance underscores need for ‘permissionless innovation’ in U.S. InsideAIPolicy.com · Charlie Mitchell
- Decoded: What is Kimi K3, an open AI model challenging OpenAI and Anthropic Business Standard · Aashish Kumar Shrivastava
- What is Kimi K3? Why this new Chinese AI model is making headlines Business Today
- Moonshot AI Unveils Kimi K3: A 2.8 Trillion-Parameter Model Challenging US AI Giants Blockonomi · Trader Edge
- New open-weight AI from China is toppling the best of OpenAI and Claude Fable Digital Trends · Sudhanshu Kumar Mangalam
- Kimi's open model K3 nears GPT-5.6 Sol and Fable 5 while signaling the end of super cheap Chinese AI The Decoder · Matthias Bastian
- Chinese startup launches “Kimi K3,” the world's largest open-weights AI model, approaching the performance of Fable and GPT-5.6 The Hans India · Kahekashan
- What is Kimi K3: Moonshot's frontier AI model that matches Claude Opus 4.8 and GPT 5.5 Digit · Vyom Ramani
- What is Kimi K3? China's new AI model aces Claude Fable 5 and GPT-5.6 Financial Express · Amritanshu Mukherjee
- Moonshot AI unveils world's largest open-source AI model as China narrows gap with US rivals South China Morning Post
- Kimi K3 shows open AI models have finally caught up with proprietary US-based rivals Neowin · Pradeep Viswanathan
- Kimi K3, and what we can still learn from the pelican benchmark Simon Willison's Weblog · Simon Willison
- Moonshot's Kimi K3 pushes Chinese AI into Fable-level territory Fortune · Nicholas Gordon
- China Just Dropped Another Bomb on America's Frontier AI Companies Gizmodo · Ece Yildirim
- [AINews] Kimi K3 2.8T-A50B: the largest open model ever released; Opus 4.8-class at Sonnet 5 pricing Latent.Space
- China's open weight Kimi K3 model snaps at the heels of US labs The Stack · Noah Bovenizer
- China's Moonshot throws down the gauntlet with Kimi K3, the world's largest open-weights model SiliconANGLE · Mike Wheatley
- Moonshot AI Releases Kimi K3: A 2.8 Trillion Parameter Open MoE Model With Kimi Delta Attention and 1M Context MarkTechPost · Asif Razzaq
- Moonshot AI releases Kimi K3, a 2.8T-parameter AI model that it says rivals Claude Opus 4.8 and GPT-5.5, and plans to release its full model weights by July 27 Kimi
- Moonshot debuts 2.8 trillion-parameter Kimi K3 Tech in Asia · Naomi Li Gan
- Moonshot AI launches Kimi K3 Constellation Research · Larry Dignan
- Chinese AI start-up Moonshot launches model challenging Anthropic's lead Financial Times
- Open-weight models now match frontier cyber performance from just four months ago at a fraction of the cost The Decoder · Matthias Bastian
- Moonshot's Kimi Upends Conventional Wisdom on US Lead Over China Bloomberg
- Kimi K3 shocked the world. These other AI models could be next Axios · Herb Scribner
- AI race splits in two as China wages open-weight insurgency Axios
- What to Know About the Chinese AI Models Rattling U.S. Stocks Wall Street Journal · Raffaele Huang
- Kimi K3 takes #1 on the Frontend Code Arena benchmark, surpassing Claude Fable 5, and scores 88.3 on Terminal Bench 2.1, trailing only GPT-5.6 Sol's 88.8 VentureBeat · Michael Nuñez
- Meet the Companies Shelling Out for Top AI Models Wall Street Journal · Belle Lin
- Kimi: Threat or menace? TechCrunch · Anthony Ha
Discussion
-
@aisecurityinst
@aisecurityinst
on x
On our cyber range “The Last Ones”, GLM-5.2 matches Opus 4.5, released ~7 months before it, while DeepSeek's V4-Pro falls below Sonnet 4.5, from ~7 months before it. [image]
-
@gdb
Greg Brockman
on x
GPT-5.6 Sol is the state of the art in cyber. Seeing significant results in applying it to finding and fixing novel vulnerabilities. Sign up as a defender to use it to secure your systems: https://openai.com/...
-
@scaling01
@scaling01
on x
open-models are lagging frontier models by 7 months on UK AISI's long-horizon cyber ranges specifically, GLM-5.2 is equivalent to Opus 4.5 on “The Last Ones” cyber range on narrower short horizon tasks GLM-5.2 performs comparably to Opus 4.6 [image]
-
@ramez
Ramez Naam
on x
1. OpenAI's 5.6 Sol beats Anthropic's still not fully available Mythos in this hacking evaluation. 2. The best open model, Kimi K3, will probably be quite similar to the US proprietary leaders. Everyone now has frontier hacking capabilities.
-
@aisecurityinst
@aisecurityinst
on x
These findings indicate a narrow window before today's frontier cyber capabilities may become widely accessible without safeguards. It's uncertain how the gap will evolve, but AISI will continue to track it and intends to test Kimi K3 once its weights are released.
-
@aisecurityinst
@aisecurityinst
on x
Our open weight model evaluations were largely unimpeded by safeguards. Of the two recent open models we tested, DeepSeek V4-Pro occasionally refused narrow cyber tasks, but this was easily circumvented by a small number of repeat attempts at refused tasks.
-
@aisecurityinst
@aisecurityinst
on x
On our narrow cyber tasks, GLM-5.2 matches Opus 4.6 and GPT-5.3-Codex, released 4 months before it. DeepSeek V4-Pro matches Opus 4.5, released 5 months before it.
-
r/singularity
r
on reddit
GPT-5.6 Sol outperforms Mythos 5 on AISI's cyber challenge
-
NewsMax.com
Charlie McCarthy
on x
China's New AI Threatens US Tech Lead
-
@deanwball
Dean W. Ball
on x
Some observations on Kimi: 1. It's a very good model! I don't think its performance can be explained away by distillation or anything like that. In agentic coding sessions, it seems pretty much on par with the best public models of Q1 2026. In my fairly limited use, it also s…
-
@scaling01
@scaling01
on x
MoonshotAI will overtake OpenAI and Anthropic before the end of the year or will they? at least that's what the hype kiddies on X want you to believe So let me make it falsifiable. They are saying: - China / MoonshotAI is catching up - they are catching up generally (not just [im…
-
@semianalysis_
@semianalysis_
on x
Similar to the panic over DeepSeek R1, some uneducated people think Kimi K3's use of linear attention (KDA) is bad for NVIDIA, HBM, DRAM, and networking because it has relatively lower KV-cache requirements. The opposite is true, and we explain why below. 👇️ 1/8🧵 [image]
-
@shakeelhashim
Shakeel
on x
Kimi K3 threatens to once again send Washington into a panic. But look a little closer, and the hysteria may be unwarranted. On Thursday, the AI industry got its second “DeepSeek moment” — this time courtesy of Moonshot, whose new Kimi K3 model has “erased America's AI lead,”
-
@teortaxestex
@teortaxestex
on x
Kimi K3 and GLM 5.2 are remarkable in that these labs *underreport* their model's performance on the most valuable evals in official announcements. In DeepSWE, K3 is not 2.5% behind Fable. More like 1%. And cheaper. We're a long way from “benchmaxxing” era. Update accordingly. [i…
-
@xeophon
Florian Brand
on x
1/5 the cost, less tokens and the same performance as Fable [image]
-
@deanwball
Dean W. Ball
on x
I wonder if the California attorney general or the European Union ai office will seek to make moonshot comply with their respective frontier ai regulations
-
@counternotions
Kontra
on x
Three decades ago China's largest export category was clothing. Time flies.
-
@davidsacks
David Sacks
on x
This is *exactly* what I predicted would happen. I said Chinese models would have advanced cyber capabilities within a matter of months and the only thing to do about it was to use AI-powered cyberdefense to protect our systems. Trying to gatekeep models doesn't work.
-
@emollick
Ethan Mollick
on x
Kimi is, as I have been saying, a very good model. But it is not a DeepSeek r1 moment, in that it is roughly where I would expect on the curve rather than an unexpected leap. It will be treated as a DeepSeek moment for a variety of reasons especially as more people hear about it.
-
@rsalakhu
Russ Salakhutdinov
on x
Congratulations to Zhilin Yang, founder and CEO of @Kimi_Moonshot, on the latest Kimi release. What a huge win for the open-source community! It feels like just yesterday Zhilin was graduating from my lab at CMU, jointly co-advised with William Cohen. Not only did he complete [im…
-
@emollick
Ethan Mollick
on x
A lot of swift conclusions are being drawn about Kimi K3 based on fairly saturated benchmarks and ELOs, rather than actually testing it on very hard problems. The AI frontier has already moved so far that a good model that is a still months behind looks like the future to many.
-
@ctjlewis
Lewis
on x
Count the number of competitors YC has put up to Kimi or DeepSeek and then let me know if you find hawkish immigration policy across three administrations to be the real problem here.
-
@zephyr_z9
@zephyr_z9
on x
After the K3 drop, Huawei has now released a real 950 SuperPoD 1 EFLOPS fp8 & 2 EFLOPS fp4 256TB of unified memory [image]
-
@negligible_cap
@negligible_cap
on x
The fact that $BABA owns 36% of Moonshot (is their biggest single investor) but dropped 4% overnight seems like evidence that there's a lot more than just Kimi to this tech selloff News on Moonshot's most recent valuation was in June, when they were looking to raise at a $30B
-
@agupta
Ankit Gupta
on x
FYI that America's moronic visa policy is almost certainly a factor in why Kimi/Moonshot is a Chinese startup and not an American one. The fact that we don't staple a green card to every AI PhD completed in America is stupid. would be more logical to seize their passports and
-
@shakeelhashim
Shakeel
on x
This is an inaccurate and irresponsible headline from Axios. [image]
-
@enzo_gte
Enzo
on x
I've been using Kimi K3 for ~16 hours now. The model is clearly good at a lot of different things (especially frontend), but non obvious reason why people are enjoying it so much is that it clearly does not follow the same rules in terms of safeguards and copyright. Kimi will h…
-
@modestproposal1
@modestproposal1
on x
The conversation around k3 is evidence of something I've been joking about. The fundamental microeconomic questions today are no different than 2 years ago, with 2 big differences: 1) there is now macro reflexivity given increasing capital market dependency 2) $120B in ARR
-
@levie
Aaron Levie
on x
This post is key. The cheaper AI gets, the more opportunity there is for the entire ecosystem - especially including end-customers - to benefit. Everything is bottlenecked by being able to successfully and cost effectively deploy AI in real workloads. Any time we can lower the…
-
@teortaxestex
@teortaxestex
on x
K3 is a startling vindication of Wenfeng's seemingly naive, idealistic thesis that culture is the best moat. Moonshot has nothing else. Compute? xAI. Pretraining science, interp, user data? Ant. RL? OAI. Ph.Ds, general data? GDM. Labor? Meta... Barely noticeable moats. Trivial. […
-
@chamath
Chamath Palihapitiya
on x
Gavin is right. This is positive for everyone except the closed frontier labs.
-
@samanthaladuc
Samantha LaDuc
on x
A bullish Kimi K3 argument: “Anything that lowers margins and increases competition at the model layer is good for every other AI layer: power, semiconductors, hyperscalers, neoclouds and yes even software.”
-
@kantrowitz
Alex Kantrowitz
on x
Weird to see the progress from Kimi K3, Grok, and Meta, and have Google absent from party
-
@modestproposal1
@modestproposal1
on x
Is it a Christensen world of modularity and models being good enough? Is it a Brian Arthur world of increasing returns. Do models commoditize? Do they “oligopolize” like IAAS did? Is capital a moat? Go back to when Llama 3 was getting rave reviews and look at the debates.
-
@omarsar0
Elvis
on x
That missing token efficiency is coming. Can't say more now, but efficient frontier long context reasoning/understanding and other TTC breakthroughs are on the horizon. Architectural improvements are often ignored in these benchmarks, but I expect major shifts by EOY.
-
@willccbb
Will Brown
on x
oh no what if the big labs can't easily recoup all of their capex on massive data centers and some of the buildout capacity gets offloaded to other providers who deliver it to the long tail of enterprises with a software stack for serving and continually improving open models
-
@liangsays
Brent Liang
on x
one of the best takes on this site re kimi. we are chronicling a mini sputnik moment right now and living through history
-
@pstasiatech
Paul Triolo
on x
...the hypothesis that they have much more advanced model checkpoints internally that are already being used for RSI. In the latter scenario, reaching RSI even a few months ahead of other labs might be enough to cement a permanent lead.
-
@limitingthe
@limitingthe
on x
“Anything that lowers margins and increases competition at the model layer is good for every other AI layer: power, semiconductors, hyperscalers, neoclouds and yes even software.”
-
@notthreadguy
@notthreadguy
on x
genuinely one of the best bull posts I've ever read > China open source ai is better than US? Long all capex beneficiaries > China open source ai is NOT better than US? Long all capex beneficiaries
-
@bare_birk
Birk
on x
Interesting post from @GavinSBaker! Kimi K3 is bad for Anthropic and OpenAI but good for all other companies. Margins will go from the frontier labs, to all other companies in the sector. Infrastructure will still be very important in both scenarios (Opensource vs closed source)
-
@shanumathew93
Shanu Mathew
on x
Great insights. Gavin nails it. If an oligopoloy of labs sustain 90% inference margins, they capture most of the economics and eventually vertically integrate the stack & squeeze chips, power, data centers, cloud and software. More competition at the model layer —> lower model…
-
@dzambhalahodl
Steven Lubka
on x
Kimi3 is a good model, but it is not a “cheap” model nor is it a “small” model. Kimi3 showed that China can train a decent competing model less than 6 months behind the US. It did not show that they can compete with Frontier at a similar efficiency advantage to Deepseek etc.
-
@davidsacks
David Sacks
on x
This is concerning. For the first time, a Chinese model Kimi K3 has taken #1 on the Frontend Code Arena and is scoring at or near the frontier on other benchmarks. Meanwhile America is tying itself in knots: politicians and bureaucrats are banning new data centers, piling on
-
@arena
@arena
on x
Big news: Kimi-K3 by @Kimi_Moonshot is now #1 in the Frontend Code Arena with 1679 pts, surpassing Claude Fable 5. This is a 17-place jump from Kimi-k2.6 (#18 -> #1). In Frontend, Kimi-K3 ranked #1 in 6 of 7 domains: Brand & Marketing, Reference-Based Design, Data & Analytics, [i…
-
@artificialanlys
@artificialanlys
on x
Kimi K3 scores 57 on the Artificial Analysis Intelligence Index. Its intelligence is comparable to Opus 4.8 and GPT-5.5 but remains behind Fable 5 and GPT-5.6 Sol. Moonshot AI has expressed plans to release the 2.8T parameter model's weights, which would make it the leading open …
-
@quxiaoyin
Xiaoyin Qu
on x
What Kimi K3 means for USA AI: 1. FYI Kimi K3 is open weight and will be released on July 27, 2026. 2. When the best open weight model exceeds the best closed-source model, how does @AnthropicAI justify its Fable pricing? Why would anyone pay for that? Not to mention the Mythos
-
@datacurve
@datacurve
on x
Kimi K3 debuts at #3 on DeepSWE. It's the first open-weights model that delivers frontier-level performance, achieving results similar to Claude Fable and GPT-5.6 Sol. [video]
-
@theo
@theo
on x
Kimi k3 is an incredible model. It is not an incredible value. In most tasks, it comes out to roughly the same cost as GPT-5.6 Sol. K3 is half the price of 5.6 Sol per token. GPT-5.6 uses half as many tokens. Price evens out. GPT-5.6 is 2x faster TPS, so it gets work done ~4x [im…
-
@aftfuture
@aftfuture
on x
.@DavidSacks is right. America is not going to regulate, restrict, and delay its way to victory in the AI race. Blocking data centers, piling on state mandates, and requiring federal preapproval for frontier models would not make us safer. It would make China stronger.
-
@rajaxg
Raja Koduri
on x
(K)impressive! As an early proof of concept, Kimi K3 designed a chip to serve a nano model built on its own architecture. In a single 48-hour autonomous run, K3 built, optimized, and verified the chip using open-source EDA tools on the Nangate 45nm library. Within 4 mm², the
-
@semianalysis_
@semianalysis_
on x
CHINA'S KIMI K3 HAS SURPASSED ALL AMERICAN MODELS IN FRONT-END CODING WHILE BEING SMALLER THAN MOST CLOSED-SOURCE FRONTIER MODELS. Great work by the @Kimi_Moonshot team. [image]
-
@scmallaby
Sebastian Mallaby
on x
I get it. But Kimi K3 also tells us that Mythos-level cyber capability is going to be freely downloadable soon. So we will go from a world in which almost nobody was getting Mythos to a world in which everyone gets it. @davidsacks, what's your story on how we address this?
-
@benbajarin
Ben Bajarin
on x
Exactly right. You can argue it's bad for the frontier model labs but all lower token costs does is drastically increase demand for compute. It's the hyperscalers who benefit from this more than anyone and even if THEY were they only ones spending, there will still be more dema…
-
@miles_brundage
Miles Brundage
on x
K3 analysis wen Also, reminder to Americans - we could have this kind of state capacity at home. Let's properly fund and unmuzzle CAISI! https://x.com/...
-
@signulll
@signulll
on x
a while ago there was a device called the palm pre, pretty revolutionary at the time. it ran something entirely new called webOS, which was built using html5. it had multi tasking, card based navigation, was very cool & slick. palm needed hundreds of engineers, years of
-
@aisecurityinst
@aisecurityinst
on x
Our first public analysis of the open/closed weight gap in frontier cyber capabilities finds it is 4-7 months with GLM-5.2 and DeepSeek V4-Pro, narrowing from 6-10 months through most of 2025. Advanced capabilities are reaching less safeguarded open models faster than before. 🧵 […
-
@crystalsssup
Crystal
on x
Kimi K3 makes me feel like one prompt is enough to build a game. You can build a game console, plug in a cartridge, and play a game on it. I just played one of my fav game “Ace Attorney” on my computer! [video]
-
@obrien
Chris O'Brien
on x
In retrospect, was it perhaps a mistake to elect a guy whose primary concern was getting eaten by sharks while on an electric boat to lead the country through this period?
-
@jukan05
Jukan
on x
Anyone loudly touting how much cheaper Kimi K3 is than Western LLMs is deliberately ignoring how much more expensive it has become compared with previous Chinese LLMs.
-
@nic_carter
Nic Carter
on x
From an ecological perspective, western frontier models are a common pool resource that Chinese open weight models exploit. They are finite, because if you exploit them too hard, they lose the incentive and resources to train new models, and everyone loses. (Chinese models aren't
-
@semianalysis_
@semianalysis_
on x
Kimi K3 2.8T is so large that it will not fit on a single NVIDIA DGX B200, even at FP4. A GB300 NVL72, B300, or MI355X system is required, as each GPU has 288 GB of memory. One optimization that could make Kimi K3 fit on B200 is to gang multiple nodes together and use a [image]
-
@zck
Zak Kukoff
on x
A few thoughts about what might happen as open models continue to achieve near-frontier performance: 1/ Closed-lab revenue is basically a function of distance between frontier outperformance and open model commoditization - call it six months. So the AI race accelerates
-
@zephyr_z9
@zephyr_z9
on x
Damn Kimi recreated Windows XP
-
@ryangreenblatt
Ryan Greenblatt
on x
Kimi K3 was significantly but not massively above my expectations. I'd tentatively guess it's similar in overall usefulness/usability to Opus 4.8 and in overall capability somewhat above Opus 4.8 (while also being somewhat more benchmaxxed). As a pretrain, it's probably somewhere
-
@nicky_sap
@nicky_sap
on x
Everyone raving about Kimi K3 so I had to give it a shot. I signed up for free and threw it a pretty nebulous prompt: 'look up Nick Saponaro and create a groundbreaking website using animation, cursor reaction, masterful design, and unique user experience. not just a parallax [vi…
-
@laschuk
@laschuk
on x
I put Kimi K3 up against Fable 5 in my benchmark. The test: clone Apple's homepage. Single prompt, one shot, zero help. Kimi K3: $0.44 (at left) Fable 5: $0.94 (at right) Same test. Same prompt. Half the price. Fable runs $10/$50 per MTok, the most expensive tokens on the [video]
-
@bhavani_00007
@bhavani_00007
on x
I tested Kimi K3 vs Claude Opus 4.8 Same prompt, an armory bay with lighting, props, and detail. Top is Kimi K3, bottom is Opus 4.8. It's not even close. Kimi K3 built a full scene with textures, proper lighting, ammo crates, weapon racks, working detail everywhere. Opus 4.8 [vid…
-
@jiahanjimliu
Jim Liu
on x
Neoclouds: The Kimi K3 Scare Kimi K3 caused a large scare in the AI trade as this Chinese open source model matched frontier models on benchmarks. Let me unpack what's actually going on. Chinese Labs have much less GPUs than American Labs and yet are able to train “just as
-
@alexfinn
Alex Finn
on x
I was wrong. I said we were a year away from Fable 5 on our desk. That day is today An open model BETTER than Fable 5 in some benchmarks just dropped Better than ChatGPT 5.6 on FrontierSWE. Better than Fable 5 on Automation Bench This fundamentally changes the AI race forever
-
@pimdewitte
Pim de Witte
on x
kimi is truly a revolutionary model [image]
-
@mweinbach
Max Weinbach
on x
It finally finished, here's the final output. Used 60% of my monthly Kimi usage on it https://macos27.kimi.page/
-
@mweinbach
Max Weinbach
on x
I asked a Kimi K3 Max agent swarm to recreate macOS 27 with real Liquid Glass and native apps in web browser and it's been going for 3 hours [image]
-
@arena
@arena
on x
Kimi-K3 just topped the Frontend Code Arena with a 76% pairwise win rate. When its output was compared head-to-head against other models on the same task, it was picked as the better output 76% of the time on average. For reference: Claude Fable 5 (63%), GPT-5.6 Sol (58%). 50% [i…
-
@arena
@arena
on x
In the Text Arena, Kimi-K3 by @Kimi_Moonshot landed #9, with 1486 pts. This is another significant improvement from Kimi-k2.6 (#38 -> #9). - Top 10 in Creative Writing, Coding and Instruction Following - #1 in three occupations: Physical & Social Science, Legal & Government, [ima…
-
@deredleritt3r
Prinz
on x
The most interesting question about Kimi K3 is whether it poses cyber risk. Kimi K3 benchmarks do not include a CyberGym score. Waiting for @AISecurityInst to bench this model.
-
@yzhang_cs
Yu Zhang
on x
K3 has now crossed the 1M context-length barrier, and DeepSeek's sparse attn has done the same. But what architecture will take us to 5M, 10M, or even longer? I'd always argue that fixed-state linear attn, especially GDN/KDA, is highly competitive here. Hybrid designs are
-
@yzhang_cs
Yu Zhang
on x
funny that K3 is great at making 3Blue1Brown videos, and the Quantile Balancing example in the blog took just a few shots to produce. https://kimi.com/...
-
@arena
@arena
on x
Kimi K3 from @Kimi_Moonshot has moved the Pareto Frontier for Code Arena: Frontend. [image]
-
Georg Zoeller
Georg Zoeller
on linkedin
The reason this all ends up with war one day is because to Americans everything everyone in the world does is about them and if they are not winning, it's always an attack. …
-
Dave Schroeder, PhD
Dave Schroeder, PhD
on linkedin
This is intentional, by design, and strategically intended in every dimension to elicit exactly this panicked reaction, to destabilize US AI development …
-
@carlquintanilla
Carl Quintanilla
on bluesky
“.. People are worried that if US companies start using Chinese models more and Anthropic less, then Anthropic will invest less. That means those US firms will lower the capex and in the end chip demand will be affected.” — @bloomberg.com — www.bloomberg.com/news/article... …
-
@metacurity.com
Cynthia Brumfield
on bluesky
China's Moonshot just issued its new Kimi K3 model, which the company said rivals the strongest offerings from OpenAI and Anthropic PBC, and this is freaking everyone out, even though it was utterly predictable. — www.bloomberg.com/news/article...
-
@carnage4life
Dare Obasanjo
on bluesky
More coverage of Kimi K3's achievements below
-
r/technology
r
on reddit
China's Moonshot unveils world's largest open AI model, closing in on US rivals
-
r/neoliberal
r
on reddit
China's open-weight Kimi model stuns AI world with frontier-level results
-
r/StockMarket
r
on reddit
China's open-weight Kimi model stuns AI world with frontier-level results
-
r/technology
r
on reddit
China's Latest A.I. Breakthrough Threatens America's Lead
-
@deanbaker13
Dean Baker
on bluesky
This doesn't look good for the Silicon Valley AI boys www.nytimes.com/2026/07/17/b... Who is going to pay trillions for their crap when China has better stuff for much less?
-
@gavinsbaker
Gavin Baker
on x
Models like Kimi K3, Grok 4.5, and Muse 1.1 may prevent the dominance of 2-3 frontier labs with 90% inference margins from hurting other AI ecosystem layers
-
@dan_jeffries1
Daniel Jeffries
on x
TLDR, not a single prediction of the AI safety folks has been correct and actual more mundane, everyday problems in the real world are all completely missed (and fixable with basic rules and engineering and time) while we continue to focus on the big fake imaginary problems of
-
@joshua_saxe
Joshua Saxe
on x
AI safety from GPT-2 to Kimi K3 Imagine a village with a wizard who one day emerges from his cave with the following tale for his fellow villagers: “Long poor, I will lead our village to prosperity by producing a series of ever more powerful magic wands as long as we're willing
-
@multiplay3r
Peter Johnston
on x
Moonshot AI, the startup behind Kimi K3 has just 300 employees? 🤯