/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

Chinese startup Moonshot releases Kimi K2 Thinking, an open-weight model it claims beats GPT-5 in agentic capabilities; source: the model cost $4.6M to train

Chinese startup Moonshot on Thursday released its latest generative artificial intelligence model which claims to beat OpenAI's ChatGPT in …

CNBC Evelyn Cheng

Context & Ripple Effects

Moonshot had already positioned Kimi K2 around a large mixture-of-experts design and benchmark performance in its earlier K2 release. K2 Thinking extends that positioning from general benchmark comparisons toward agentic work, while retaining an open-weight distribution model.

The reported $4.6 million training cost makes the release relevant not only as a performance claim but as a challenge to the assumption that frontier-style agent capabilities require only the largest closed-model budgets. The GPT-5 comparison remains Moonshot’s claim, rather than an independently established result.

First-order effects

  • Developers and enterprises can evaluate and deploy an open-weight alternative aimed at agentic workloads, rather than relying exclusively on proprietary model APIs.
  • Moonshot gains a sharper competitive pitch: claimed agent performance against OpenAI alongside a disclosed, comparatively modest training-cost figure.

Second-order effects

  • Open-weight and proprietary model vendors face added pressure to demonstrate agent reliability on comparable tasks, not simply publish broad benchmark results.
  • If users can reproduce the reported capability, buyers gain leverage in model selection and deployment negotiations because an open-weight option can be assessed for self-hosted use.

Third-order effects

  • The release points toward competition shifting from raw model scale to cost-adjusted performance on multi-step tasks, where reproducibility and operational reliability will determine whether vendor claims translate into adoption.
  • A sustained stream of capable open-weight agent models could broaden buyer choice and reduce dependence on a small set of hosted-model providers, though this depends on real-world deployment results and support ecosystems.

The trend: Agentic AI competition is increasingly becoming a contest over useful task performance and deployment economics, not just parameter counts or closed-model access.

Discussion

  • @kimi_moonshot @kimi_moonshot on x
    🚀 Hello, Kimi K2 Thinking! The Open-Source Thinking Agent Model is here. 🔹 SOTA on HLE (44.9%) and BrowseComp (60.2%) 🔹 Executes up to 200 - 300 sequential tool calls without human interference 🔹 Excels in reasoning, agentic search, and coding 🔹 256K context window Built [image]
  • @bubblebabyboi Bubble Boi on x
    > Be Chinese > Take everybody's innovations > Improve it 10x > Give it away for free? why do they keep doing this? :
  • @awnihannun Awni Hannun on x
    The new 1 Trillion parameter Kimi K2 Thinking model runs well on 2 M3 Ultras in its native format - no loss in quality! The model was quantization aware trained (qat) at int4. Here it generated ~3500 tokens at 15 toks/sec using pipeline-parallelism in mlx-lm: [video]
  • @emollick Ethan Mollick on x
    Testing Kimi K-2 has reminded me of how insane it is that firms picking AIs are treating them as fungible based on benchmarks Kimi & Grok & Claude & every other model have strengths, quirks & weaknesses that can make a big difference in aggregate Develop your own benchmarks!
  • @pvncher Eric Provencher on x
    Ran the new Kimi K2 Thinking on Repo Bench, via @OpenRouterAI Results are pretty disappointing. Not only was the score generally very low (much worse than Kimi 0905), running this bench was noticeably slower than pretty much any recently tested model. [image]
  • @eliebakouch Elie on x
    the score are insane, very cool to see native int4 quantization for the MoE layers > To overcome this challenge, we adopt Quantization-Aware Training (QAT) during the post-training phase, applying INT4 weight-only quantization to the MoE components. It allows K2 Thinking to
  • @industrlpolicy Dhiraj on x
    Each successive state-of-the-art AI model that Chinese labs release as open source is becoming more efficient than its predecessor ie lower training costs and higher efficiency inference with performance. https://www.cnbc.com/... [image]
  • @tim_dettmers Tim Dettmers on x
    I guess we are now very close to open-weights vs closed source parity. Can't test it since I am traveling and my laptop broke (😭), but many people say it's better than Sonnet/Gemini/Grok. Very exciting times!
  • @ofirpress Ofir Press on x
    Congrats to the Kimi K2 team on the great numbers on our SWE-bench Verified, SWE-bench Multilingual and SciCode benchmarks!! [image]
  • @tphuang @tphuang on x
    It is really a good thing when open source AI lab stops just comparing themselves to other open src models, but compare to the best closed src ones. I will have to test out Kimi K2, but this looks very promising.
  • @zephyr_z9 @zephyr_z9 on x
    10x cheaper than GPT-5 and 20x cheaper than Sonnet 4.5 No wonder OpenAI needs the government money
  • @artificialanlys @artificialanlys on x
    MoonshotAI has released Kimi K2 Thinking, a new reasoning variant of Kimi K2 that achieves #1 in the Tau2 Bench Telecom agentic benchmark and is potentially the new leading open weights model Kimi K2 Thinking is one of the largest open weights models ever, at 1T total parameters …
  • @teortaxestex @teortaxestex on x
    > All benchmark results are reported under INT4 precision. Do you understand what a flex this was. They go toe to toe with GPT-5 on the heaviest, longest-range tasks, with hundreds of tool calls. ALL IN INT4. «Convert to fp8 if you need» Frontier lab. [image]
  • @clementdelangue Clem on x
    The AI frontier is open-source!
  • @simonw Simon Willison on x
    Kimi K2 Thinking is “potentially the new leading open weight” according to @ArtificialAnlys
  • @ai_for_success AshutoshShrivastava on x
    What Moonshot AI is doing is unbelievable. HLE and IMO both reached SOTA, and it's SOTA worldwide, not only SOTA in China. This company's valuation (Moonshot AI) is only 0.5% of OpenAI's, 2% of Anthropic's and Grok's. And they manage to bring up some model like this.
  • @deedydas Deedy on x
    🚨 Today is a turning point in AI. A Chinese open source model is #1. Kimi K2 Thinking scored 51% in Humanity's Last Exam, higher than GPT-5 and every other model. $0.6/M in, $2.5/M output. The best at writing, and does 15tps on two Mac M3 Ultras! Seminal moment in AI. Try it [ima…
  • @felixclc_ @felixclc_ on x
    Funniest part about K2 thinking is that it's native Int4. For those unawares, Int-4 was Ampere exclusive within the nvidia ecosystem. The only way to run the model at its native precision is using two generation old hardware 😂
  • @philipkiely Philip Kiely on x
    Kimi K2 was already the best model for creative writing. K2 Thinking takes it to the next level for deep research and technical content. I tested Kimi head-to-head with GPT 5 Pro on highly technical writing with thousands of tokens of context + agentic search. Results: [video]
  • @crystalsssup Crystal on x
    It's SOTA, not only open weights SOTA :)
  • @eliebakouch Elie on x
    > “200-300 sequential tool calls” this is really the impressive part of this release imo, can't wait to see how they did it [image]
  • @awnihannun Awni Hannun on x
    I'm a bit giddy over the fact that this is by all visible measures a frontier level model, if not THE frontier model, for agentic tasks. And you can run it. In it's native precision. On 2 M3 Ultras. Pretty fast. In MLX.
  • @thom_wolf Thomas Wolf on x
    Is this another DeepSeek moment? Open-source passing closed-source again Should we expect this every couple months now?
  • @eliebakouch Elie on x
    we're very close to 50% on HLE, and bonus point: it's with an open model :) [image]
  • @andrew_n_carr Andrew Carr on x
    70% on SWE bench verified 30% terminal bench those are two intuitive thresholds for “actually useful and not frustrating” coding assistant. Kimi k2 thinking got 71.3% on SWE-Bench Verified 47.1% on Terminal-Bench
  • @natolambert Nathan Lambert on x
    Thoughts on Kimi K2 Thinking Congrats to the Moonshot AI team on the awesome open release. For close followers of Chinese AI models, this isn't shocking, but more inflection points are coming. Pressure is building on US labs with more expensive models. https://www.interconnects.a…
  • @emostaque Emad on x
    Congratulations to @Kimi_Moonshot for achieving state of the art on many benchmarks & open sourcing the model! The gap between closed & open continues to narrow even as the cost of increasingly economically valuable tokens collapses K2 has its own unique vibe too, try it out!
  • @emostaque Emad on x
    Necessity is the mother of invention Also - training optimally on small amounts of chips with focus on data means the Chinese models take 10-100x less compute to run as well & have that cost advantage $150/mGPT 4.5 vs $0.5/m DeepSeek v3 etc
  • @gnotuy Yutong on x
    Today, we're releasing Kimi K2 Thinking, our best open-source model. What makes it different isn't just the benchmarks, though it achieves SOTA results on Humanity's Last Exam, BrowseComp, and other challenging tests. What matters is how it thinks. It reminds me of the minds on
  • @_mira___mira_ Mira on x
    > 200-300 sequential tool calls You guys know base models, right? If they optimize for “immediate expected value”, then reasoning models optimize for “trajectory expected value”. I think there's sort of a “kelly criterion for language models” at the distribution level. If an
  • @rasbt Sebastian Raschka on x
    Exciting big Kimi K2 Thinking release! More experts, fewer heads, and even more thinking! [image]
  • @casper_hansen_ Casper Hansen on x
    K2 Thinking released with Heavy Mode! K2 Thinking Heavy Mode employs an efficient parallel strategy: it first rolls out eight trajectories simultaneously, then reflectively aggregates all outputs to generate the final result. BETTER than gpt-5-pro at HLE :) [image]
  • @andrewcurran_ Andrew Curran on x
    News thread for Thursday, Nov 6th. Looks like multiple announcements today. Kimi K2-Thinking is live in Kimi-chat now, not in app for me yet. Benchmarks are leaking ahead of the official announcement. [image]
  • @skylermiao7 Skyler Miao on x
    Congrats! Great to see an OSS model outperforming closed ones on HLE, and another model adopting interleaved thinking with impressive results. Really cool to see our efforts with partners, like @OpenRouterAI and @cline, to support interleaved thinking now benefiting more users
  • @yulun_du Yulun Du on x
    🧠💡Now we give you an open-sourced frontier thinking agent model! Enjoy!
  • @scaling01 @scaling01 on x
    The gap is closing. China is catching up. Kimi-K2 Thinking crushes GPT-5 and Claude 4.5 Sonnet in several benchmarks, while costing 6 times less compared to Sonnet It's the best open-source model period Its core focus is on agentic tasks and software development. It can now [imag…
  • @burkov @burkov on x
    A new open-weight Kimi K2 Thinking claims to be comparable to GPT-5 and Sonnet 4.5. It's a Mixture-of-Experts (MoE) model with a total of 1T parameters and 32B activated parameters for token generation. The context length is 256K, which makes it competitive for coding. Weights [i…
  • @kalomaze @kalomaze on x
    the date is november 6th, 2025. you can download the most powerful agentic artificial intelligence in the world for free, under a permissive license and so it is a good day!
  • @emostaque Emad on x
    A note on costs/compute Base Kimi K2 model used 2.8m H800 hours with 14.8 trillion tokens, about $5.6m worth Details of post training for reasoning not given, but it is likely max 20% more (excluding data prep!) Would be < $3m for sota if they had Blackwell chip access
  • @deanbaker13 Dean Baker on bluesky
    China's progress with open-source models doesn't look like good news for the Silicon Valley AI boys venturebeat.com/ai/moonshots...
  • @markriedl Mark Riedl on bluesky
    The Chinese Kimi K2 thinking model beats GPT and Claude on some benchmarks.  This analysis from @natolambert.bsky.social is a good overview iew of what is going on www.interconnects.ai/p/kimi-k2- th...
  • @ErikJonker@mastodon.social Erik Jonker on mastodon
    Reading thoughts about a new Chinese openweights AI model, Kimi K2 Thinking.  —  https://www.interconnects.ai/ ...  https://simonwillison.net/...  #ai #kimiK2
  • r/singularity r on reddit
    KIMI K2 Thinking Benchmarks
  • r/LocalLLaMA r on reddit
    Kimi K2 Thinking and DeepSeek R1 Architectures Side by Side
  • r/LocalLLaMA r on reddit
    My Hands-On Review of Kimi K2 Thinking: The Open-Source AI That's Changing the Game
  • r/singularity r on reddit
    The chinese did it, KIMI K2 surpassed GPT-5.
  • r/LocalLLaMA r on reddit
    Kimi released Kimi K2 Thinking, an open-source trillion-parameter reasoning model