Sources: ByteDance is pretraining an AI model with up to 10T parameters, roughly 3x larger than Kimi K3 and larger than the 8T estimate for Anthropic's Mythos 5
TikTok owner training a model three times larger than Moonshot's Kimi K3 — ByteDance is training an AI model that could approach …
Financial Times
Context & Ripple Effects
ByteDance's reported pretraining effort follows its earlier plan to build an AI model primarily on Huawei Ascend 910B chips and arrives weeks after Moonshot released the 2.8T-parameter Kimi K3. The comparison makes training scale a more visible competitive marker among Chinese model developers.
First-order effects
ByteDance becomes the reported leader in this coverage on nominal model size, with a training run of up to 10T parameters versus Kimi K3's 2.8T parameters.
Moonshot's Kimi K3 now serves as the immediate domestic scale benchmark ByteDance is seeking to exceed; the report does not establish comparative model performance.
Second-order effects
ByteDance's earlier Ascend-chip plan gains strategic weight because a frontier-scale pretraining run concentrates demand on the hardware and capacity available to the company.
Moonshot and Tencent face a more explicit scale comparison in their model positioning, after Tencent's Hy3-preview was reported at 295B parameters.
Third-order effects
If larger pretraining runs continue to define frontier competition, Chinese AI labs will compete not only on model releases but on sustained access to training infrastructure and capacity allocation.
The pattern shifts attention from parameter counts alone toward whether model developers can convert large training runs into deployable, competitive systems.
The trend: Chinese AI developers are escalating frontier-model competition through larger pretraining commitments, making compute access and capacity planning central strategic differentiators.
holy fck it's a 10T model.. that's the same size as mythos... the ramifications of china open sourcing this has not been thought through enough. if you thought the openai hugging face attack was bad, just wait till a swarm of chinese agents descend upon your company database [ima…
ByteDance is reportedly training a model with up to 10tn parameters. Anthropic's Mythos 5 is estimated at ~8tn. Fable 5 at ~5tn. The US strategy is to make frontier AI harder/unreachable for China. China keeps going up at the frontier anyway? …
first 5T, now it's 10T if this is true and they finish training a 10T model this year, then a lot of people (including me) have been horribly wrong I was predicting that a 10T chinese model wouldn't happen until early to mid 2027
Bytedance has undefeated champion of video gen AI, Seedance. Now they are cooking 10T model which is equivalent to Mythos' size. I don't know how they would serve this massive model without B300, but they will find the way.
ByteDance is going for a 5T model (im sorry I included google in my AGI tier list half a year ago. i was blinded by the hopium, and had longer timelines before Mythos) [image]
🚨 Reports indicate ByteDance is discussing training a massive LLM model with over 5 trillion parameters Looks like China is REALLY scaling up now. For reference, Kimi K3 is “only” 2.8 trillion params ByteDance's founder is also supposedly against distilling western models [image]
registering a prediction that this is going to be a flop the one weakness that particularly competent vp had was not knowing (nor being interested in knowing) any of the details in pretraining i hope i'm wrong!