/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

Sources: Moonshot is seeking access to more Nvidia Blackwell chips to prepare for Kimi K4's development, after training K3 on Nvidia chips, including Blackwell

Beijing-based startup Moonshot AI, flush from the breakout success of its recently released model, Kimi K3 …

The Information

Context & Ripple Effects

Moonshot’s move follows a rapid Kimi release cadence: it introduced an open-weight Kimi K2.6 model focused on long-horizon coding and then prepared K3 as a substantially larger model. The reported Blackwell use for K3 makes additional access a concrete input to the next development cycle.

The story matters because it ties Moonshot’s model roadmap directly to Nvidia capacity. Nvidia has positioned GB200 Blackwell systems as particularly suited to MoE workloads, an architecture Moonshot had already used in earlier Kimi models.

First-order effects

  • Moonshot is seeking more Blackwell capacity for Kimi K4, making access to Nvidia hardware a near-term constraint on the startup’s next training program.
  • Nvidia gains another prospective demand source from a lab that reportedly trained K3 on its chips, including Blackwell.

Second-order effects

  • Moonshot’s need for additional capacity raises the value of securing scarce advanced-compute access relative to simply publishing model weights or benchmark claims.
  • Competing model developers pursuing larger or more capable successors face added pressure to lock in comparable training infrastructure; suppliers serving AI clusters benefit if this demand translates into deployments.

Third-order effects

  • If leading labs repeatedly pair faster model iterations with larger accelerator allocations, compute access becomes a more durable differentiator in frontier-model development rather than a one-time procurement decision.
  • The pattern reinforces an AI infrastructure bottleneck in which model competition is increasingly shaped by the availability and allocation of advanced hardware.

The trend: Frontier AI development is becoming more tightly coupled to access to high-end accelerator capacity, especially for labs advancing successive large-model generations.

Discussion

  • @kimi_moonshot @kimi_moonshot on x
    Releasing the model weights and technical report of Kimi K3.  Kimi K3 is our most capable model …
  • @natolambert Nathan Lambert on x
    Kimi K3 license. It's inspired by MIT but distinctly non-commercial, where any company making over $20M/yr must get a specific commercial deal (and display Kimi K3 if over 100M users or $20M/mo revenue) [image]
  • @ollama @ollama on x
    Kimi K3 is now available on Ollama's cloud. To use it with Claude Code, run: ollama launch claude —model kimi-k3:cloud Currently Kimi K3 requires a Pro or Max subscription, and consumes extra usage credits. We're quickly working on adding capacity to expand access.
  • @suchenzang Susan Zhang on x
    theme of k3 design: numerical stability (signal prop, precision, etc) at scale is still a Hard Problem a short🧵while reading through...
  • @arena @arena on x
    Big update: Among open-weight models, Kimi K3 (Max) is #1 in the Agent Arena with +9.75% net-improvement …
  • @notsurajgaud Suraj Gaud on x
    if k3 disappears tomorrow, i got you boys. [image]
  • @sentdex Harrison Kinsley on x
    Promises made, promises kept, Kimi K3 is now officially open weights. The license looks good and fair. Big day for open models.
  • @adxtyahq Aditya on x
    Reminder: you only need a $35,000/month GPU budget, or a one-time $500,000 investment, to avoid a $99/month Kimi K3 subscription.
  • @hsvsphere @hsvsphere on x
    Based based based based based based based Do not fall for Big Labs that make Frontier Intelligence look …
  • @teknium @teknium on x
    Kimi made it to Hugging Face and is now open weights, with a paper too!
  • @yuchenj_uw Yuchen Jin on x
    Kimi K3 weights are open now. It is the largest open-weight model so far (2.8T parameters). It's only slightly behind Claude Fable 5 Max and GPT-5.6 Sol Max. Pretty excited about this model, and again, a Muon victory! [image]
  • @totheagi Ning on x
    we got the full Kimi K3, 2.8T params, running on 80x RTX 5090s. 20 tok/s single stream, day one, untuned. …
  • @emollick Ethan Mollick on x
    Been reading the weights of the most powerful open weights Ai model yet released. Page one starts strong: mostly negatives, a late run of zeros, and one unexpectedly large positive value. [image]
  • @yzhang_cs Yu Zhang on x
    lost for words rn.
  • @zephyr_z9 @zephyr_z9 on x
    104B Active, much bigger than I thought
  • @kimi_moonshot @kimi_moonshot on x
    We've open-sourced FlashKDA, our high-performance CUTLASS-based implementation of Kimi Delta Attention kernels. It delivers 1.72×-2.22× prefill speedup over the flash-linear-attention baseline on H20, and works as a drop-in backend for flash-linear-attention. Explore on GitHub: h…
  • @kevinsxu Kevin S. Xu on x
    Here is how I see the Kimi K3 license works in real life: - Clause 2, the Model as a Service clause …
  • @eliebakouch Elie on x
    this scaling law is a piece of art, kimi K3 recipe improves by ~2.5x over kimi K2 recipe the tech report is amazing [image]
  • @rauchg Guillermo Rauch on x
    Kimi K3 is the most powerful open-weight model in the world. It's now available on @vercel AI Gateway from 🇺🇸 USA-based inference providers at high availability & performance. We've signed ZDR (Zero-Data Retention) agreements, which you can enable for all your token traffic.
  • @cognition @cognition on x
    Kimi K3 is now available in Devin Desktop and CLI. On FrontierCode 1.1, Kimi K3 is the first open source model we tested that approaches frontier-level performance. [image]
  • @petergostev Peter Gostev on x
    Kimi K3 License: MIT + these 2 points: 1) Large AI hosting companies earning over $20M/year need a separate agreement. 2) Products over 100M users or $20M/month revenue must display “Kimi K3.”
  • @modal @modal on x
    Kimi K3 is live on Modal. Moonshot has shipped the world's first open 3T-class model, and we're a day zero launch partner. We trained a custom DFlash speculator for K3's novel architecture so you can run it faster, losslessly. The most capable open model we've worked with by far.
  • @perrymetzger Perry E. Metzger on x
    You can download a SOTA AI model and, if you have enough hardware, just run it, customize it …
  • @peterom @peterom on x
    I'm very happy that Kimi are going w/ a license that will let them benefit from large companies that make money off their models 3 years ago, I was a license purist but capitalism and openness must work together & this approach - similar to LTX's - feels like a nice balance
  • @nathanbenaich Nathan Benaich on x
    America being a fast follower to Chinese frontier open AI with all inference platforms hosting Kimi K3 this morning is a fun vibe shift.
  • @suchenzang Susan Zhang on x
    there are 3 axes that issues come across when Scaling, and we can roughly bin them into: 1) sequence length (context window/attention) 2) depth (# layers) 3) width (MoE) each one comes with particular sets of bandages to constrain dynamic range (1/n) [image]
  • @teortaxestex @teortaxestex on x
    Kimi K3 tech report does not disclose the number of training tokens A pity, because this is straight up the biggest model trained in China ever, and I suspect it's well over 2e25 FLOPs, probably 3-4e25. The report is overall great. [image]
  • @zephyr_z9 @zephyr_z9 on x
    Very impressive stuff from Kimi Training efficiency improved by 2.5x [image]
  • @clementdelangue Clem on x
    Kimi K3 has been released on HF and already top 1 trending with 4,000+ likes in 30 mins. Fastest release growth ever so far!! [image]
  • @__tinygrad__ @__tinygrad__ on x
    This is a great sustainable business model for open weights. It's free if you run it yourself, but if you are running a cloud providing it to others for money, you should have to share profits. (from Kimi K3 License) [image]
  • @zackkorman Zack Korman on x
    Moonshot AI also had some problems with Kimi K3 agents trying to escape, so they just made a better sandbox. Funny how that works. The sandbox environment they run is even open source. Love the transparency here. [image]
  • @eliebakouch Elie on x
    wow with kimi K3 open weights, moonshot is also open sourcing a significative improvement over deepseek deepEP which is currently used by everyone in the open (vllm/sglang), this is huge https://github.com/... [image]
  • @fireworksai_hq @fireworksai_hq on x
    Kimi K3 is live on Fireworks. Day 0, inference and training. US-hosted, and zero data retention. This is the first frontier open model in the 3 trillion parameter class. It sports 1M context, native vision, and reasoning that rivals the top closed models. Boom. [image]
  • @presidentlin @presidentlin on x
    They updated the license it's no longer modified-mit but a kimi-k3 license.  I think they found a nice balance tldr 2. …
  • @nebiustf @nebiustf on x
    Kimi K3 is now available on Token Factory.  We're excited to announce that Nebius Token Factory …
  • r/GeminiAI r on reddit
    How is Google getting outpaced by Kimi-K3?  A tech monolith shouldn't be losing like this.
  • r/ClaudeCode r on reddit
    Kimi K3 has become open-weights just as of a few minutes ago.
  • r/Anthropic r on reddit
    Anthropic makes money from its API customers, right?  Well, Kimi K3's weighs went out a few minutes ago.  They're in trouble.
  • @kyleichan Kyle Chan on x
    Kimi K3 partly trained on Blackwells along with Qwen3.8 and DeepSeek, according to The Information. https://www.theinformation.com/ ...
  • @leomschwartz Leo Schwartz on x
    Important reporting from @theinformation today on how Moonshot trained Kimi K3 using Nvidia chips and its plans for K4 https://www.theinformation.com/ ... [image]
  • @sayreevan Evan Swarztrauber on x
    According to this report, Kimi K3 trained on export controlled NVIDIA Blackwell chips. The open vs. closed debate should not distract us from chip smuggling and lax enforcement. One can be pro-open-source AI while wanting to prevent advanced chips from going to China.