/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

← → days · ↑ ↓ browse · Enter similar · o open

A study by Meta researchers suggests that training LLMs to predict multiple tokens at once, instead of just the next token, results in better and faster models

LLM approach to predict multiple tokens KAN: Kolmogorov-Arnold Networks —"promising alternatives to Multi-Layer Perceptrons" [image] Ethan / @ethan_smith_20 : it was only briefly touched upon, but is it correct that multi-token prediction is only valid in the case of greedy decoding? also IMHO it seems like it'd make more sense and parameter efficient to have 1 shared head but then 1-2 branching transformer layers for each token. [image] @lchoshen : Pretrain to predict the future At each step the model predicts n-tokens Performance: 😃 Inference time: ✖️3 Training time: same @AIatMeta @FabianGloeckle @byoubii @b_roziere @dfpazr @syhw https://arxiv.org/... [image] Yoshinari Fujinuma / @akkikiki : I got to read the multi-token prediction paper in depth today, and IMO, the most important takeaway is “Multi-token prediction improves performance and unlocks efficient byte level training”. Simple idea, interesting results & takeaways. https://arxiv.org/... [image] Benjamin Lefaudeux / @bentheegg : Multi token prediction that works, really nice paper which I think will be foundational https://export.arxiv.org/... (1/N) Woojin Kim / @woojinrad : Better & Faster Large Language Models via Multi-token Prediction 🤔 Interesting work—training language models to predict multiple future tokens at once. They show that models trained with 4-token prediction are up to 3x faster at inference. #LLMs https://arxiv.org/... [image] Benjamin Lefaudeux / @bentheegg : It's not a completely new idea, even recently (see Medusa, attached) but a huge difference here is that it covers training and inference, it's not a post-hoc change (via fine tuning). Arguably having multi token as part of the training is a bigger deal [image] Chenwei Cui / @ccui42 : There's a reason why people use next-1-token prediction not next-N-token prediction: the latter requires conditional independence within the N-tokens. Curious how this paper resolves this. Observation: 1. N-token prediction but with overlapping of N-1 tokens. 2. Looks pretty... Rohan Paul / @rohanpaul_ai : Meta's groundbreaking paper - “Better & Faster Large Language Models via Multi-token Prediction” ✨ Models trained with 4-token prediction are up to 3 times faster at inference, even with large batch sizes. 🔥 📌 Large language models such as GPT and Llama are trained with a... [image] Matthew Leavitt / @leavittron : Great thread from @teortaxesTex on the Multi-token Prediction paper ( https://arxiv.org/...)! One thing I will add: the authors propose that multi-token prediction may work because it implicitly weights tokens by relevance. I'd love to see how multi-token prediction composes... [image] @cto_junior : One of the important observations in Multi-token prediction paper is not that it's fast (otherwise it would be a bummer) but that prediction n > 1 tokens also improves model's accuracy on multiple coding evals [image] Leo Boytsov / @srchvrs : “More specifically, at each position in the training corpus, we ask the model to predict the following n tokens using n independent output heads, operating on top of a shared model trunk. Considering multi-token prediction as an auxiliary training task, we measure improved... @_akhaliq : Meta announces Better & Faster Large Language Models via Multi-token Prediction Large language models such as GPT and Llama are trained with a next-token prediction loss. In this work, we suggest that training language models to predict multiple future tokens at [image] Amaury Hayat / @amaury_hayat : Really cool paper by @FabianGloeckle @byoubii, @b_roziere, David Lopez-Paz and @syhw ! I've heard about it several times over the last few months, and I'm glad to see it out :) Yuchen Jin / @yuchenj_uw : Will Multi-token prediction revolutionize LLMs, which currently rely on next-token prediction? Interesting section: [image] Leo Boytsov / @srchvrs : Again, I eagerly reposted this paper and several people expressed this exact concern. It did not help on some tasks. Alex Dimakis / @alexgdimakis : This Medusa architecture (predicting multiple future tokens with different heads on the same branch) improves performance on generative tasks, it seems. Unfortunately the benefits come at 10B scale or more so we would never see them by running smaller experiments. Also... Vik / @vikhyatk : this is neat, might also be useful to apply on MoE MLP layers. [image] Aran Komatsuzaki / @arankomatsuzaki : Btw this multiple token training is not a panacea.  The performance gain depends on the target task.  It leads to no perf gain or slight degradation on some multiple choice questions.  It leads to minior improvement on summarization and little to no improvement on arithmetic tasks. @mrcatid : After reading the new Meta “Multi-token Prediction” paper, I want to try adding some FFN heads to predict the next 4 tokens or 8 bytes. Looks like they didn't try that but seems like it's a nice design optimization point. Mira / @_mira___mira_ : There's obvious-seeming latency improvements I was wondering why nobody was doing: * Multiple tokens at once * Ability to exit early from the transformer stack * Use a cheaper computation to get “top 10 most likely” and speculatively compute all of them while the larger runs @teortaxestex : On the futility of small scaling experiments; or why frontier labs will out-innovate you Devastating figure [image] Aran Komatsuzaki / @arankomatsuzaki : Meta presents Better & Faster Large Language Models via Multi-token Prediction - training language models to predict multiple future tokens at once results in higher sample efficiency - up to 3x faster at inference https://arxiv.org/... [image] Yam Peleg / @yampeleg : Tried this already, it didn't improve the training for me. Probably is more efficient above certain amount of tokens that I didn't reach. Cool idea and very easy to try! Joey / @shxf0072 : tldr: > make n<=4 prediction lm head, > better perf & faster speculative decoding > with no overhead [image] @teortaxestex : Good question re Multi-Token Prediction. My answer: «at this model scale, 4 is optimal». And maybe at this data scale. It probably stacks well with curricula. Just as we're increasing seqlen throughout training, we may gradually increase text complexity *and* lookahead window. [image] Cody Blakeney / @code_star : I'm really excited about ideas like this, but before people get too worked up you should know this seems to be a domain specific intervention. Thats *ok* though. This might be a very useful piece of making code models or adapting models to be code models. I'm also willing to... [image] Tianle Cai / @tianle_cai : Wow, Medusa can be used for pre-training and leads to a better and faster generation! 😍 @teortaxestex : A detail I overlooked in Multi-Token Prediction: they test induction capability emergence using a synthetic dataset of children stories and small models, showing positive results until 30M. Wonder what dataset they mean, hm. Guess I have some intuition. [image] Archie / @archiexzzz : good read: better & faster llms via multi-token prediction by meta. training language models to predict multiple future tokens at once. https://arxiv.org/... [image] Saurabh Kumar / @drummatick : @archiexzzz Skipgram architecture did something similar with word2vec I haven't followed the development beyond that, but my guess is beyond n=5 the performance might be dropping and gains might not be uniform even for n>2 Kaizhao Liang / @kyleliang5 : old but interesting idea, this is the first time I see it working in the wild. https://arxiv.org/... what matters is the “intention”. If LLM wants to simulate human thought process, it needs t o be able to “know” what it wants to generate before it generates it. A mental draft... Ethan Mollick / @emollick : This may end up being a big deal: Usually LLMs just predict the next token in a sequence, one at a time, but if you have them predict the next several tokens at once you get significantly better performance, faster, and with no added costs. The gains are better for bigger models [image] LinkedIn: Jack FitzGerald : Instead of only training with next token prediction, some model parameters are instead distributed to n 1-layer heads on top of a shared trunk of n-1 layers … Shang-You T. : New paper “Better & Faster Large Language Models via Multi-token Prediction” can change direction of LLMs. … Forums: Hacker News : Better and Faster Large Language Models via Multi-Token Prediction

VentureBeat Ben Dickson

Context & Ripple Effects

Meta’s work arrives alongside its broader Llama push: it had detailed Llama 3’s 8B- and 70B-parameter releases weeks earlier, while this research targets the training objective rather than simply scaling model size. The proposed design keeps a shared trunk and adds independent output heads for several future tokens.

The idea moved beyond a paper shortly afterward, when Meta released research models using multi-token prediction. That progression makes the study a useful signal about an alternative route to improving model quality and training efficiency.

First-order effects

  • Meta’s researchers gain a training recipe that, according to the study, improves performance and sample efficiency by supervising several future tokens at each position.
  • Model builders experimenting with the method must alter the output layer and training loss to support multiple independent prediction heads, while retaining a shared model trunk.

Second-order effects

  • If the reported gains reproduce across model families, teams face a trade-off between conventional next-token training and more complex objectives that may extract more learning from the same training data.
  • The later release of pre-trained multi-token models gives researchers a concrete artifact for testing whether the paper’s training gains transfer to downstream use cases.

Third-order effects

  • The work points toward model training becoming a more important optimization surface alongside parameter count: architectural and objective-level changes could matter as much as additional scale.
  • Whether multi-token objectives become standard will depend on whether their training benefits persist under varied decoding and deployment conditions, rather than on benchmark gains alone.

The trend: LLM developers are increasingly seeking efficiency and quality gains through training objectives and architectures, not only larger models and more data.

Discussion

  • @omarsar0 Elvis on x
    The most exciting LLM paper of the week was the one from Gloeckle et al. that aims to train better and faster LLM via multi-token prediction.  It's an impressive research paper so I had lots of thoughts as usual, especially because it attempts to push LLMs forward...
  • @intrstllrninja @intrstllrninja on x
    Multi-token prediction is 3x faster using self-speculative decoding while also improving performance on tasks like coding and algorithmic reasoning as it emphasizes on longer-term dependencies [image]
  • @intuitmachine Carlos E. Perez on x
    1/n Beyond Next-Word Prediction: Multi-Token Prediction Imagine you're trying to learn a new language. You could start by memorizing individual words, but that would only get you so far. To truly understand the language, you need to grasp how words connect and form meaningful... …
  • @itshesamsheikh Hesam on x
    Two very important papers just released in ML community on my reading list: Better & Faster Large Language Models via Multi-token Prediction by @Meta — LLM approach to predict multiple tokens KAN: Kolmogorov-Arnold Networks —"promising alternatives to Multi-Layer Perceptrons" [im…
  • @ethan_smith_20 Ethan on x
    it was only briefly touched upon, but is it correct that multi-token prediction is only valid in the case of greedy decoding? also IMHO it seems like it'd make more sense and parameter efficient to have 1 shared head but then 1-2 branching transformer layers for each token. [imag…
  • @lchoshen @lchoshen on x
    Pretrain to predict the future At each step the model predicts n-tokens Performance: 😃 Inference time: ✖️3 Training time: same @AIatMeta @FabianGloeckle @byoubii @b_roziere @dfpazr @syhw https://arxiv.org/... [image]
  • @akkikiki Yoshinari Fujinuma on x
    I got to read the multi-token prediction paper in depth today, and IMO, the most important takeaway is “Multi-token prediction improves performance and unlocks efficient byte level training”. Simple idea, interesting results & takeaways. https://arxiv.org/... [image]
  • @bentheegg Benjamin Lefaudeux on x
    Multi token prediction that works, really nice paper which I think will be foundational https://export.arxiv.org/... (1/N)
  • @woojinrad Woojin Kim on x
    Better & Faster Large Language Models via Multi-token Prediction 🤔 Interesting work—training language models to predict multiple future tokens at once. They show that models trained with 4-token prediction are up to 3x faster at inference. #LLMs https://arxiv.org/... [image]
  • @bentheegg Benjamin Lefaudeux on x
    It's not a completely new idea, even recently (see Medusa, attached) but a huge difference here is that it covers training and inference, it's not a post-hoc change (via fine tuning). Arguably having multi token as part of the training is a bigger deal [image]
  • @ccui42 Chenwei Cui on x
    There's a reason why people use next-1-token prediction not next-N-token prediction: the latter requires conditional independence within the N-tokens. Curious how this paper resolves this. Observation: 1. N-token prediction but with overlapping of N-1 tokens. 2. Looks pretty...
  • @rohanpaul_ai Rohan Paul on x
    Meta's groundbreaking paper - “Better & Faster Large Language Models via Multi-token Prediction” ✨ Models trained with 4-token prediction are up to 3 times faster at inference, even with large batch sizes. 🔥 📌 Large language models such as GPT and Llama are trained with a... [ima…
  • @leavittron Matthew Leavitt on x
    Great thread from @teortaxesTex on the Multi-token Prediction paper ( https://arxiv.org/...)! One thing I will add: the authors propose that multi-token prediction may work because it implicitly weights tokens by relevance. I'd love to see how multi-token prediction composes... […
  • @cto_junior @cto_junior on x
    One of the important observations in Multi-token prediction paper is not that it's fast (otherwise it would be a bummer) but that prediction n > 1 tokens also improves model's accuracy on multiple coding evals [image]
  • @srchvrs Leo Boytsov on x
    “More specifically, at each position in the training corpus, we ask the model to predict the following n tokens using n independent output heads, operating on top of a shared model trunk. Considering multi-token prediction as an auxiliary training task, we measure improved...
  • @_akhaliq @_akhaliq on x
    Meta announces Better & Faster Large Language Models via Multi-token Prediction Large language models such as GPT and Llama are trained with a next-token prediction loss. In this work, we suggest that training language models to predict multiple future tokens at [image]
  • @amaury_hayat Amaury Hayat on x
    Really cool paper by @FabianGloeckle @byoubii, @b_roziere, David Lopez-Paz and @syhw ! I've heard about it several times over the last few months, and I'm glad to see it out :)
  • @yuchenj_uw Yuchen Jin on x
    Will Multi-token prediction revolutionize LLMs, which currently rely on next-token prediction? Interesting section: [image]
  • @srchvrs Leo Boytsov on x
    Again, I eagerly reposted this paper and several people expressed this exact concern. It did not help on some tasks.
  • @alexgdimakis Alex Dimakis on x
    This Medusa architecture (predicting multiple future tokens with different heads on the same branch) improves performance on generative tasks, it seems. Unfortunately the benefits come at 10B scale or more so we would never see them by running smaller experiments. Also...
  • @vikhyatk Vik on x
    this is neat, might also be useful to apply on MoE MLP layers. [image]
  • @arankomatsuzaki Aran Komatsuzaki on x
    Btw this multiple token training is not a panacea.  The performance gain depends on the target task.  It leads to no perf gain or slight degradation on some multiple choice questions.  It leads to minior improvement on summarization and little to no improvement on arithmetic task…
  • @mrcatid @mrcatid on x
    After reading the new Meta “Multi-token Prediction” paper, I want to try adding some FFN heads to predict the next 4 tokens or 8 bytes. Looks like they didn't try that but seems like it's a nice design optimization point.
  • @_mira___mira_ Mira on x
    There's obvious-seeming latency improvements I was wondering why nobody was doing: * Multiple tokens at once * Ability to exit early from the transformer stack * Use a cheaper computation to get “top 10 most likely” and speculatively compute all of them while the larger runs
  • @teortaxestex @teortaxestex on x
    On the futility of small scaling experiments; or why frontier labs will out-innovate you Devastating figure [image]
  • @arankomatsuzaki Aran Komatsuzaki on x
    Meta presents Better & Faster Large Language Models via Multi-token Prediction - training language models to predict multiple future tokens at once results in higher sample efficiency - up to 3x faster at inference https://arxiv.org/... [image]
  • @yampeleg Yam Peleg on x
    Tried this already, it didn't improve the training for me. Probably is more efficient above certain amount of tokens that I didn't reach. Cool idea and very easy to try!
  • @shxf0072 Joey on x
    tldr: > make n<=4 prediction lm head, > better perf & faster speculative decoding > with no overhead [image]
  • @teortaxestex @teortaxestex on x
    Good question re Multi-Token Prediction. My answer: «at this model scale, 4 is optimal». And maybe at this data scale. It probably stacks well with curricula. Just as we're increasing seqlen throughout training, we may gradually increase text complexity *and* lookahead window. [i…
  • @code_star Cody Blakeney on x
    I'm really excited about ideas like this, but before people get too worked up you should know this seems to be a domain specific intervention. Thats *ok* though. This might be a very useful piece of making code models or adapting models to be code models. I'm also willing to... […
  • @tianle_cai Tianle Cai on x
    Wow, Medusa can be used for pre-training and leads to a better and faster generation! 😍
  • @teortaxestex @teortaxestex on x
    A detail I overlooked in Multi-Token Prediction: they test induction capability emergence using a synthetic dataset of children stories and small models, showing positive results until 30M. Wonder what dataset they mean, hm. Guess I have some intuition. [image]
  • @archiexzzz Archie on x
    good read: better & faster llms via multi-token prediction by meta. training language models to predict multiple future tokens at once. https://arxiv.org/... [image]
  • @drummatick Saurabh Kumar on x
    @archiexzzz Skipgram architecture did something similar with word2vec I haven't followed the development beyond that, but my guess is beyond n=5 the performance might be dropping and gains might not be uniform even for n>2
  • @kyleliang5 Kaizhao Liang on x
    old but interesting idea, this is the first time I see it working in the wild. https://arxiv.org/... what matters is the “intention”. If LLM wants to simulate human thought process, it needs t o be able to “know” what it wants to generate before it generates it. A mental draft...
  • @emollick Ethan Mollick on x
    This may end up being a big deal: Usually LLMs just predict the next token in a sequence, one at a time, but if you have them predict the next several tokens at once you get significantly better performance, faster, and with no added costs. The gains are better for bigger models …