Models like Kimi K3, Grok 4.5, and Muse 1.1 may prevent the dominance of 2-3 frontier labs with 90% inference margins from hurting other AI ecosystem layers
Related coverage puts Kimi K3 in a mixed position: it has posted strong coding-benchmark results, while a separate assessment argues that this does not make it a frontier model or evidence of a broader geopolitical lead.
The immediate business backdrop is rising customer scrutiny of AI spending and the prospect that OpenAI, Meta, and SpaceXAI compete more explicitly on cost efficiency. The significance of Kimi, Grok, and Muse is therefore less any single benchmark than whether credible substitutes constrain the economics of the leading providers.
First-order effects
Kimi K3, Grok 4.5, and Muse 1.1 expand the set of models buyers can evaluate for high-value workloads, reducing the practical dependence on OpenAI and Anthropic where their performance is sufficient.
The availability of more credible alternatives increases pressure on frontier labs to defend inference pricing, performance, and enterprise value rather than rely solely on scarcity.
Second-order effects
Enterprise customers gain more leverage to route workloads by cost and task fit, reinforcing the cost-efficiency competition already visible in coverage of OpenAI, Meta, and SpaceXAI.
Model providers that cannot sustain a clear capability or distribution advantage may have to compete through lower-cost inference, specialized performance, or easier deployment, rather than premium general-purpose access alone.
Third-order effects
If several labs can remain credible on commercially important tasks, AI value capture may be distributed more broadly across models, application builders, and infrastructure rather than concentrating in two or three high-margin inference suppliers.
That outcome is not established by benchmark results alone: it depends on whether challengers can translate point performance into reliable, broadly deployable products and sustained customer adoption.
The trend: The story is one data point in a shift from frontier-model scarcity toward a more contested AI market where usable substitutes constrain model-provider pricing power.
I've been using Kimi K3 for ~16 hours now. The model is clearly good at a lot of different things (especially frontend), but non obvious reason why people are enjoying it so much is that it clearly does not follow the same rules in terms of safeguards and copyright. Kimi will h…
The conversation around k3 is evidence of something I've been joking about. The fundamental microeconomic questions today are no different than 2 years ago, with 2 big differences: 1) there is now macro reflexivity given increasing capital market dependency 2) $120B in ARR
This post is key. The cheaper AI gets, the more opportunity there is for the entire ecosystem - especially including end-customers - to benefit. Everything is bottlenecked by being able to successfully and cost effectively deploy AI in real workloads. Any time we can lower the…
K3 is a startling vindication of Wenfeng's seemingly naive, idealistic thesis that culture is the best moat. Moonshot has nothing else. Compute? xAI. Pretraining science, interp, user data? Ant. RL? OAI. Ph.Ds, general data? GDM. Labor? Meta... Barely noticeable moats. Trivial. […
A bullish Kimi K3 argument: “Anything that lowers margins and increases competition at the model layer is good for every other AI layer: power, semiconductors, hyperscalers, neoclouds and yes even software.”
Is it a Christensen world of modularity and models being good enough? Is it a Brian Arthur world of increasing returns. Do models commoditize? Do they “oligopolize” like IAAS did? Is capital a moat? Go back to when Llama 3 was getting rave reviews and look at the debates.
That missing token efficiency is coming. Can't say more now, but efficient frontier long context reasoning/understanding and other TTC breakthroughs are on the horizon. Architectural improvements are often ignored in these benchmarks, but I expect major shifts by EOY.
oh no what if the big labs can't easily recoup all of their capex on massive data centers and some of the buildout capacity gets offloaded to other providers who deliver it to the long tail of enterprises with a software stack for serving and continually improving open models
...the hypothesis that they have much more advanced model checkpoints internally that are already being used for RSI. In the latter scenario, reaching RSI even a few months ahead of other labs might be enough to cement a permanent lead.
“Anything that lowers margins and increases competition at the model layer is good for every other AI layer: power, semiconductors, hyperscalers, neoclouds and yes even software.”
genuinely one of the best bull posts I've ever read > China open source ai is better than US? Long all capex beneficiaries > China open source ai is NOT better than US? Long all capex beneficiaries
Interesting post from @GavinSBaker! Kimi K3 is bad for Anthropic and OpenAI but good for all other companies. Margins will go from the frontier labs, to all other companies in the sector. Infrastructure will still be very important in both scenarios (Opensource vs closed source)
Great insights. Gavin nails it. If an oligopoloy of labs sustain 90% inference margins, they capture most of the economics and eventually vertically integrate the stack & squeeze chips, power, data centers, cloud and software. More competition at the model layer —> lower model…
Kimi3 is a good model, but it is not a “cheap” model nor is it a “small” model. Kimi3 showed that China can train a decent competing model less than 6 months behind the US. It did not show that they can compete with Frontier at a similar efficiency advantage to Deepseek etc.
Kimi k3 is an incredible model. It is not an incredible value. In most tasks, it comes out to roughly the same cost as GPT-5.6 Sol. K3 is half the price of 5.6 Sol per token. GPT-5.6 uses half as many tokens. Price evens out. GPT-5.6 is 2x faster TPS, so it gets work done ~4x [im…
(K)impressive! As an early proof of concept, Kimi K3 designed a chip to serve a nano model built on its own architecture. In a single 48-hour autonomous run, K3 built, optimized, and verified the chip using open-source EDA tools on the Nangate 45nm library. Within 4 mm², the
Exactly right. You can argue it's bad for the frontier model labs but all lower token costs does is drastically increase demand for compute. It's the hyperscalers who benefit from this more than anyone and even if THEY were they only ones spending, there will still be more dema…