Chinese startup Moonshot releases Kimi K2 Thinking, an open-weight model it claims beats GPT-5 in agentic capabilities; source: the model cost $4.6M to train
Chinese startup Moonshot on Thursday released its latest generative artificial intelligence model which claims to beat OpenAI's ChatGPT in …
CNBC Evelyn Cheng
Context & Ripple Effects
Moonshot had already positioned Kimi K2 around a large mixture-of-experts design and benchmark performance in its earlier K2 release. K2 Thinking extends that positioning from general benchmark comparisons toward agentic work, while retaining an open-weight distribution model.
The reported $4.6 million training cost makes the release relevant not only as a performance claim but as a challenge to the assumption that frontier-style agent capabilities require only the largest closed-model budgets. The GPT-5 comparison remains Moonshot’s claim, rather than an independently established result.
First-order effects
- Developers and enterprises can evaluate and deploy an open-weight alternative aimed at agentic workloads, rather than relying exclusively on proprietary model APIs.
- Moonshot gains a sharper competitive pitch: claimed agent performance against OpenAI alongside a disclosed, comparatively modest training-cost figure.
Second-order effects
- Open-weight and proprietary model vendors face added pressure to demonstrate agent reliability on comparable tasks, not simply publish broad benchmark results.
- If users can reproduce the reported capability, buyers gain leverage in model selection and deployment negotiations because an open-weight option can be assessed for self-hosted use.
Third-order effects
- The release points toward competition shifting from raw model scale to cost-adjusted performance on multi-step tasks, where reproducibility and operational reliability will determine whether vendor claims translate into adoption.
- A sustained stream of capable open-weight agent models could broaden buyer choice and reduce dependence on a small set of hosted-model providers, though this depends on real-world deployment results and support ecosystems.
The trend: Agentic AI competition is increasingly becoming a contest over useful task performance and deployment economics, not just parameter counts or closed-model access.
Related: AI cost per useful task · Model buyer power · Moonshot · Kimi K2 Thinking · Moonshot's earlier Kimi K2 architecture and benchmark results
Related Coverage
- 5 Thoughts on Kimi K2 Thinking Interconnects · Nathan Lambert
- Moonshot's Kimi K2 Thinking emerges as leading open source AI, outperforming GPT-5, Claude Sonnet 4.5 on key benchmarks VentureBeat · Carl Franzen
- Kimi K2 Thinking Simon Willison's Weblog · Simon Willison
- Moonshot's $4.6 million ‘Kimi K2 Thinking’ takes top spots on reasoning benchmarks Implicator.ai · Maria Garcia
- OpenAI's leaked GPT-5.1 ‘Thinking’ model could outsmart Gemini 3 Pro — here's why that matters Tom's Guide · Amanda Caswell
- China's open-source AI closes the gap The Rundown AI
- Kimi K2 Thinking Crushes GPT-5, Claude 4.5 Sonnet in Key Benchmarks Analytics India Magazine · Siddharth Jindal
- Kimi K2 Thinking, a SOTA open-source trillion-parameter reasoning model Hacker News
Discussion
-
@kimi_moonshot
@kimi_moonshot
on x
🚀 Hello, Kimi K2 Thinking! The Open-Source Thinking Agent Model is here. 🔹 SOTA on HLE (44.9%) and BrowseComp (60.2%) 🔹 Executes up to 200 - 300 sequential tool calls without human interference 🔹 Excels in reasoning, agentic search, and coding 🔹 256K context window Built [image]
-
@bubblebabyboi
Bubble Boi
on x
> Be Chinese > Take everybody's innovations > Improve it 10x > Give it away for free? why do they keep doing this? :
-
@awnihannun
Awni Hannun
on x
The new 1 Trillion parameter Kimi K2 Thinking model runs well on 2 M3 Ultras in its native format - no loss in quality! The model was quantization aware trained (qat) at int4. Here it generated ~3500 tokens at 15 toks/sec using pipeline-parallelism in mlx-lm: [video]
-
@emollick
Ethan Mollick
on x
Testing Kimi K-2 has reminded me of how insane it is that firms picking AIs are treating them as fungible based on benchmarks Kimi & Grok & Claude & every other model have strengths, quirks & weaknesses that can make a big difference in aggregate Develop your own benchmarks!
-
@pvncher
Eric Provencher
on x
Ran the new Kimi K2 Thinking on Repo Bench, via @OpenRouterAI Results are pretty disappointing. Not only was the score generally very low (much worse than Kimi 0905), running this bench was noticeably slower than pretty much any recently tested model. [image]
-
@eliebakouch
Elie
on x
the score are insane, very cool to see native int4 quantization for the MoE layers > To overcome this challenge, we adopt Quantization-Aware Training (QAT) during the post-training phase, applying INT4 weight-only quantization to the MoE components. It allows K2 Thinking to
-
@industrlpolicy
Dhiraj
on x
Each successive state-of-the-art AI model that Chinese labs release as open source is becoming more efficient than its predecessor ie lower training costs and higher efficiency inference with performance. https://www.cnbc.com/... [image]
-
@tim_dettmers
Tim Dettmers
on x
I guess we are now very close to open-weights vs closed source parity. Can't test it since I am traveling and my laptop broke (😭), but many people say it's better than Sonnet/Gemini/Grok. Very exciting times!
-
@ofirpress
Ofir Press
on x
Congrats to the Kimi K2 team on the great numbers on our SWE-bench Verified, SWE-bench Multilingual and SciCode benchmarks!! [image]
-
@tphuang
@tphuang
on x
It is really a good thing when open source AI lab stops just comparing themselves to other open src models, but compare to the best closed src ones. I will have to test out Kimi K2, but this looks very promising.
-
@zephyr_z9
@zephyr_z9
on x
10x cheaper than GPT-5 and 20x cheaper than Sonnet 4.5 No wonder OpenAI needs the government money
-
@artificialanlys
@artificialanlys
on x
MoonshotAI has released Kimi K2 Thinking, a new reasoning variant of Kimi K2 that achieves #1 in the Tau2 Bench Telecom agentic benchmark and is potentially the new leading open weights model Kimi K2 Thinking is one of the largest open weights models ever, at 1T total parameters …
-
@teortaxestex
@teortaxestex
on x
> All benchmark results are reported under INT4 precision. Do you understand what a flex this was. They go toe to toe with GPT-5 on the heaviest, longest-range tasks, with hundreds of tool calls. ALL IN INT4. «Convert to fp8 if you need» Frontier lab. [image]
-
@clementdelangue
Clem
on x
The AI frontier is open-source!
-
@simonw
Simon Willison
on x
Kimi K2 Thinking is “potentially the new leading open weight” according to @ArtificialAnlys
-
@ai_for_success
AshutoshShrivastava
on x
What Moonshot AI is doing is unbelievable. HLE and IMO both reached SOTA, and it's SOTA worldwide, not only SOTA in China. This company's valuation (Moonshot AI) is only 0.5% of OpenAI's, 2% of Anthropic's and Grok's. And they manage to bring up some model like this.
-
@deedydas
Deedy
on x
🚨 Today is a turning point in AI. A Chinese open source model is #1. Kimi K2 Thinking scored 51% in Humanity's Last Exam, higher than GPT-5 and every other model. $0.6/M in, $2.5/M output. The best at writing, and does 15tps on two Mac M3 Ultras! Seminal moment in AI. Try it [ima…
-
@felixclc_
@felixclc_
on x
Funniest part about K2 thinking is that it's native Int4. For those unawares, Int-4 was Ampere exclusive within the nvidia ecosystem. The only way to run the model at its native precision is using two generation old hardware 😂
-
@philipkiely
Philip Kiely
on x
Kimi K2 was already the best model for creative writing. K2 Thinking takes it to the next level for deep research and technical content. I tested Kimi head-to-head with GPT 5 Pro on highly technical writing with thousands of tokens of context + agentic search. Results: [video]
-
@crystalsssup
Crystal
on x
It's SOTA, not only open weights SOTA :)
-
@eliebakouch
Elie
on x
> “200-300 sequential tool calls” this is really the impressive part of this release imo, can't wait to see how they did it [image]
-
@awnihannun
Awni Hannun
on x
I'm a bit giddy over the fact that this is by all visible measures a frontier level model, if not THE frontier model, for agentic tasks. And you can run it. In it's native precision. On 2 M3 Ultras. Pretty fast. In MLX.
-
@thom_wolf
Thomas Wolf
on x
Is this another DeepSeek moment? Open-source passing closed-source again Should we expect this every couple months now?
-
@eliebakouch
Elie
on x
we're very close to 50% on HLE, and bonus point: it's with an open model :) [image]
-
@andrew_n_carr
Andrew Carr
on x
70% on SWE bench verified 30% terminal bench those are two intuitive thresholds for “actually useful and not frustrating” coding assistant. Kimi k2 thinking got 71.3% on SWE-Bench Verified 47.1% on Terminal-Bench
-
@natolambert
Nathan Lambert
on x
Thoughts on Kimi K2 Thinking Congrats to the Moonshot AI team on the awesome open release. For close followers of Chinese AI models, this isn't shocking, but more inflection points are coming. Pressure is building on US labs with more expensive models. https://www.interconnects.a…
-
@emostaque
Emad
on x
Congratulations to @Kimi_Moonshot for achieving state of the art on many benchmarks & open sourcing the model! The gap between closed & open continues to narrow even as the cost of increasingly economically valuable tokens collapses K2 has its own unique vibe too, try it out!
-
@emostaque
Emad
on x
Necessity is the mother of invention Also - training optimally on small amounts of chips with focus on data means the Chinese models take 10-100x less compute to run as well & have that cost advantage $150/mGPT 4.5 vs $0.5/m DeepSeek v3 etc
-
@gnotuy
Yutong
on x
Today, we're releasing Kimi K2 Thinking, our best open-source model. What makes it different isn't just the benchmarks, though it achieves SOTA results on Humanity's Last Exam, BrowseComp, and other challenging tests. What matters is how it thinks. It reminds me of the minds on
-
@_mira___mira_
Mira
on x
> 200-300 sequential tool calls You guys know base models, right? If they optimize for “immediate expected value”, then reasoning models optimize for “trajectory expected value”. I think there's sort of a “kelly criterion for language models” at the distribution level. If an
-
@rasbt
Sebastian Raschka
on x
Exciting big Kimi K2 Thinking release! More experts, fewer heads, and even more thinking! [image]
-
@casper_hansen_
Casper Hansen
on x
K2 Thinking released with Heavy Mode! K2 Thinking Heavy Mode employs an efficient parallel strategy: it first rolls out eight trajectories simultaneously, then reflectively aggregates all outputs to generate the final result. BETTER than gpt-5-pro at HLE :) [image]
-
@andrewcurran_
Andrew Curran
on x
News thread for Thursday, Nov 6th. Looks like multiple announcements today. Kimi K2-Thinking is live in Kimi-chat now, not in app for me yet. Benchmarks are leaking ahead of the official announcement. [image]
-
@skylermiao7
Skyler Miao
on x
Congrats! Great to see an OSS model outperforming closed ones on HLE, and another model adopting interleaved thinking with impressive results. Really cool to see our efforts with partners, like @OpenRouterAI and @cline, to support interleaved thinking now benefiting more users
-
@yulun_du
Yulun Du
on x
🧠💡Now we give you an open-sourced frontier thinking agent model! Enjoy!
-
@scaling01
@scaling01
on x
The gap is closing. China is catching up. Kimi-K2 Thinking crushes GPT-5 and Claude 4.5 Sonnet in several benchmarks, while costing 6 times less compared to Sonnet It's the best open-source model period Its core focus is on agentic tasks and software development. It can now [imag…
-
@burkov
@burkov
on x
A new open-weight Kimi K2 Thinking claims to be comparable to GPT-5 and Sonnet 4.5. It's a Mixture-of-Experts (MoE) model with a total of 1T parameters and 32B activated parameters for token generation. The context length is 256K, which makes it competitive for coding. Weights [i…
-
@kalomaze
@kalomaze
on x
the date is november 6th, 2025. you can download the most powerful agentic artificial intelligence in the world for free, under a permissive license and so it is a good day!
-
@emostaque
Emad
on x
A note on costs/compute Base Kimi K2 model used 2.8m H800 hours with 14.8 trillion tokens, about $5.6m worth Details of post training for reasoning not given, but it is likely max 20% more (excluding data prep!) Would be < $3m for sota if they had Blackwell chip access
-
@deanbaker13
Dean Baker
on bluesky
China's progress with open-source models doesn't look like good news for the Silicon Valley AI boys venturebeat.com/ai/moonshots...
-
@markriedl
Mark Riedl
on bluesky
The Chinese Kimi K2 thinking model beats GPT and Claude on some benchmarks. This analysis from @natolambert.bsky.social is a good overview iew of what is going on www.interconnects.ai/p/kimi-k2- th...
-
@ErikJonker@mastodon.social
Erik Jonker
on mastodon
Reading thoughts about a new Chinese openweights AI model, Kimi K2 Thinking. — https://www.interconnects.ai/ ... https://simonwillison.net/... #ai #kimiK2
-
r/singularity
r
on reddit
KIMI K2 Thinking Benchmarks
-
r/LocalLLaMA
r
on reddit
Kimi K2 Thinking and DeepSeek R1 Architectures Side by Side
-
r/LocalLLaMA
r
on reddit
My Hands-On Review of Kimi K2 Thinking: The Open-Source AI That's Changing the Game
-
r/singularity
r
on reddit
The chinese did it, KIMI K2 surpassed GPT-5.
-
r/LocalLLaMA
r
on reddit
Kimi released Kimi K2 Thinking, an open-source trillion-parameter reasoning model