/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

Sources: ByteDance is pretraining an AI model with up to 10T parameters, roughly 3x larger than Kimi K3 and larger than the 8T estimate for Anthropic's Mythos 5

TikTok owner training a model three times larger than Moonshot's Kimi K3  —  ByteDance is training an AI model that could approach …

Financial Times

Context & Ripple Effects

ByteDance's reported pretraining effort follows its earlier plan to build an AI model primarily on Huawei Ascend 910B chips and arrives weeks after Moonshot released the 2.8T-parameter Kimi K3. The comparison makes training scale a more visible competitive marker among Chinese model developers.

First-order effects

  • ByteDance becomes the reported leader in this coverage on nominal model size, with a training run of up to 10T parameters versus Kimi K3's 2.8T parameters.
  • Moonshot's Kimi K3 now serves as the immediate domestic scale benchmark ByteDance is seeking to exceed; the report does not establish comparative model performance.

Second-order effects

  • ByteDance's earlier Ascend-chip plan gains strategic weight because a frontier-scale pretraining run concentrates demand on the hardware and capacity available to the company.
  • Moonshot and Tencent face a more explicit scale comparison in their model positioning, after Tencent's Hy3-preview was reported at 295B parameters.

Third-order effects

  • If larger pretraining runs continue to define frontier competition, Chinese AI labs will compete not only on model releases but on sustained access to training infrastructure and capacity allocation.
  • The pattern shifts attention from parameter counts alone toward whether model developers can convert large training runs into deployable, competitive systems.

The trend: Chinese AI developers are escalating frontier-model competition through larger pretraining commitments, making compute access and capacity planning central strategic differentiators.

Discussion

  • @cryptopunk7213 @cryptopunk7213 on x
    holy fck it's a 10T model.. that's the same size as mythos... the ramifications of china open sourcing this has not been thought through enough. if you thought the openai hugging face attack was bad, just wait till a swarm of chinese agents descend upon your company database [ima…
  • @jukan05 Jukan on x
    FT: BYTEDANCE IS IN THE PRE-TRAINING STAGE OF A MODEL WITH UP TO 10 TRILLION PARAMETERS, APPROACHING MYTHOS IN SCALE. [image]
  • @zephyr_z9 @zephyr_z9 on x
    They will probably distill it Serving 10T models doesn't make sense right now
  • @saridder Steve R. on bluesky
    garbage in, garbage out.
  • @LukaszOlejnik@mastodon.social Lukasz Olejnik on mastodon
    ByteDance is reportedly training a model with up to 10tn parameters.  Anthropic's Mythos 5 is estimated at ~8tn.  Fable 5 at ~5tn.  The US strategy is to make frontier AI harder/unreachable for China.  China keeps going up at the frontier anyway? …
  • @mark_k Mark Kretschmann on x
    The size of this rumored ByteDance model is doubling every day. Yesterday it was still rumored to be “5T”, today, it's “10T”... 🤔
  • @scaling01 @scaling01 on x
    first 5T, now it's 10T if this is true and they finish training a 10T model this year, then a lot of people (including me) have been horribly wrong I was predicting that a 10T chinese model wouldn't happen until early to mid 2027
  • @jun_song Jun Song on x
    Bytedance has undefeated champion of video gen AI, Seedance. Now they are cooking 10T model which is equivalent to Mythos' size. I don't know how they would serve this massive model without B300, but they will find the way.
  • @scaling01 @scaling01 on x
    ByteDance is going for a 5T model (im sorry I included google in my AGI tier list half a year ago. i was blinded by the hopium, and had longer timelines before Mythos) [image]
  • @lentils80 @lentils80 on x
    🚨 Reports indicate ByteDance is discussing training a massive LLM model with over 5 trillion parameters Looks like China is REALLY scaling up now. For reference, Kimi K3 is “only” 2.8 trillion params ByteDance's founder is also supposedly against distilling western models [image]
  • @suchenzang Susan Zhang on x
    registering a prediction that this is going to be a flop the one weakness that particularly competent vp had was not knowing (nor being interested in knowing) any of the details in pretraining i hope i'm wrong!
  • @chamath Chamath Palihapitiya on x
    And if reports are accurate, with zero distillation to jumpstart it. Thus proving many things that are more value destructive than value accreting...