Sources: Moonshot plans to launch Kimi K3, China's largest model to date with 2T-3T parameters, in the coming days; it's expected to outperform Claude Opus 4.8
Kimi K3 expected to exceed performance of Claude Opus 4.8 in sign of narrowing gap between US and China on frontier AI
Financial Times
Context & Ripple Effects
Moonshot has moved quickly through the Kimi line: K2 used a trillion-parameter mixture-of-experts design, K2 Thinking was positioned around agentic work, and K2.5 added multimodal processing. Its most recent K2.6 release emphasized long-horizon coding and agent-swarm capabilities under a modified MIT License.
That product cadence has coincided with rising commercial and financing ambitions, including reported ARR growth and large funding valuations. A new frontier-scale Kimi release therefore matters not just as a benchmark claim, but as a test of whether Moonshot can convert increasingly capable models into durable platform standing.
First-order effects
Moonshot would gain a new flagship model for competing for customers, developers, and capital; the claimed comparison with Claude Opus 4.8 will put its evaluation results under unusually close scrutiny.
Anthropic's Claude lineup becomes the immediate reference point in a high-profile China-versus-US performance comparison, though the reported advantage remains an expectation until independent testing is available.
Second-order effects
Other Chinese model providers may face pressure to answer with larger or more specialized releases, particularly in coding, tool use, multimodal work, and agent-style workflows where Moonshot has already concentrated its Kimi updates.
Enterprise buyers and developers evaluating Kimi against US frontier models gain another potential option, but practical adoption will hinge on reproducible performance and model-access terms rather than parameter scale alone.
Third-order effects
If successive Kimi releases continue to narrow reported performance gaps, frontier AI competition is likely to become less centered on a small group of US labs and more on the ability of regional providers to pair model capability with distribution, capital, and developer ecosystems.
The sequence also reinforces an unresolved industry question: whether ever-larger models remain the primary competitive signal, or whether open-weight availability and workflow-specific reliability become the more durable differentiators.
The trend: This is one data point in the globalization of frontier-model competition, as Chinese AI companies combine rapid release cycles, open-weight strategies, and larger training runs to challenge US reference models.
Chinese AI start-up Moonshot to launch model challenging Anthropic's lead * Set to release as early as tonight * 2-3T, largest Chinese model to date * Benchmark performance Opus 4.8 < K3 < Fable * Attention Residuals & Kimi Linear * Fundraising at $31.5B https://as.ft.com/...
Kimi k3 is being released tonight, via FT -2-3t parameters (Opus4.8 has about 1.5t) -1m context -Expected to exceed Opus 4.8 performance! The time when China was six months behind is over. History is presumably being made today. [image]
Kimi-K3 vs Opus-4.8 Flappybird test. Kimi is significantly better than Opus. That's the reason why I claimed it as Opus-5 level. Test done in @arena [video]
Kimi K3 is currently leading Claude Fable 5, GPT-5.6, and Grok 4.5 creating game demo. A mystery Arena .ai model called “Kivine” is reportedly Kimi K3. In this AI Game Watch, I break down the first game and 3D demos shared by creators on X. [video]
Kimi K3 just went global and beta testers are saying it matches Opus-level output at nearly half the price. 🚀 China clearly isn't slowing down in this AI race. Overhyped or the real deal? Drop your take 👇 [image]
Kimi k3 starts rolling out. Official release is imminent! Super freaking excited for the evals. Will it bear opus 4.8? Could be a real game changer, literally.
Reminder that Opus 4.8 came out at thee end of May. So frontier leadership went from 6+ months to 1-1.5 months. Unclear how much of this is because of wide distillation and how far internal models are at American labs but the direction is clear that China is catching up...
At $3/$15 in out tokens eh, like i get why it's priced what it is and the compute needed but i have no desire to pay that given it doesn't seem token efficient
Kimi K3 being 2.8T parameter and 1M context is cool but show me the sparsity, show me the price How quick can the inference providers scale this to 200 tok/s This is what I care about!!!! Efficient HUGE models
This is gorgeous! What stands out in contrast to the Claude ad is that the vibes here are entirely on possibilities, positive possibilities, and optimism.
Here are a few benchmark scores of K3 that have been officially confirmed This is a Fable/Sol class model that is strictly better than Opus 4.8 across the board at Sonnet pricing. Insane [image]