A study by Meta researchers suggests that training LLMs to predict multiple tokens at once, instead of just the next token, results in better and faster models
LLM approach to predict multiple tokens KAN: Kolmogorov-Arnold Networks —"promising alternatives to Multi-Layer Perceptrons" [image] Ethan / @ethan_smith_20 : it was only briefly touched upon, but is it correct that multi-token prediction is only valid in the case of greedy decoding? also IMHO it seems like it'd make more sense and parameter efficient to have 1 shared head but then 1-2 branching transformer layers for each token. [image] @lchoshen : Pretrain to predict the future At each step the model predicts n-tokens Performance: 😃 Inference time: ✖️3 Training time: same @AIatMeta @FabianGloeckle @byoubii @b_roziere @dfpazr @syhw https://arxiv.org/... [image] Yoshinari Fujinuma / @akkikiki : I got to read the multi-token prediction paper in depth today, and IMO, the most important takeaway is “Multi-token prediction improves performance and unlocks efficient byte level training”. Simple idea, interesting results & takeaways. https://arxiv.org/... [image] Benjamin Lefaudeux / @bentheegg : Multi token prediction that works, really nice paper which I think will be foundational https://export.arxiv.org/... (1/N) Woojin Kim / @woojinrad : Better & Faster Large Language Models via Multi-token Prediction 🤔 Interesting work—training language models to predict multiple future tokens at once. They show that models trained with 4-token prediction are up to 3x faster at inference. #LLMs https://arxiv.org/... [image] Benjamin Lefaudeux / @bentheegg : It's not a completely new idea, even recently (see Medusa, attached) but a huge difference here is that it covers training and inference, it's not a post-hoc change (via fine tuning). Arguably having multi token as part of the training is a bigger deal [image] Chenwei Cui / @ccui42 : There's a reason why people use next-1-token prediction not next-N-token prediction: the latter requires conditional independence within the N-tokens. Curious how this paper resolves this. Observation: 1. N-token prediction but with overlapping of N-1 tokens. 2. Looks pretty... Rohan Paul / @rohanpaul_ai : Meta's groundbreaking paper - “Better & Faster Large Language Models via Multi-token Prediction” ✨ Models trained with 4-token prediction are up to 3 times faster at inference, even with large batch sizes. 🔥 📌 Large language models such as GPT and Llama are trained with a... [image] Matthew Leavitt / @leavittron : Great thread from @teortaxesTex on the Multi-token Prediction paper ( https://arxiv.org/...)! One thing I will add: the authors propose that multi-token prediction may work because it implicitly weights tokens by relevance. I'd love to see how multi-token prediction composes... [image] @cto_junior : One of the important observations in Multi-token prediction paper is not that it's fast (otherwise it would be a bummer) but that prediction n > 1 tokens also improves model's accuracy on multiple coding evals [image] Leo Boytsov / @srchvrs : “More specifically, at each position in the training corpus, we ask the model to predict the following n tokens using n independent output heads, operating on top of a shared model trunk. Considering multi-token prediction as an auxiliary training task, we measure improved... @_akhaliq : Meta announces Better & Faster Large Language Models via Multi-token Prediction Large language models such as GPT and Llama are trained with a next-token prediction loss. In this work, we suggest that training language models to predict multiple future tokens at [image] Amaury Hayat / @amaury_hayat : Really cool paper by @FabianGloeckle @byoubii, @b_roziere, David Lopez-Paz and @syhw ! I've heard about it several times over the last few months, and I'm glad to see it out :) Yuchen Jin / @yuchenj_uw : Will Multi-token prediction revolutionize LLMs, which currently rely on next-token prediction? Interesting section: [image] Leo Boytsov / @srchvrs : Again, I eagerly reposted this paper and several people expressed this exact concern. It did not help on some tasks. Alex Dimakis / @alexgdimakis : This Medusa architecture (predicting multiple future tokens with different heads on the same branch) improves performance on generative tasks, it seems. Unfortunately the benefits come at 10B scale or more so we would never see them by running smaller experiments. Also... Vik / @vikhyatk : this is neat, might also be useful to apply on MoE MLP layers. [image] Aran Komatsuzaki / @arankomatsuzaki : Btw this multiple token training is not a panacea. The performance gain depends on the target task. It leads to no perf gain or slight degradation on some multiple choice questions. It leads to minior improvement on summarization and little to no improvement on arithmetic tasks. @mrcatid : After reading the new Meta “Multi-token Prediction” paper, I want to try adding some FFN heads to predict the next 4 tokens or 8 bytes. Looks like they didn't try that but seems like it's a nice design optimization point. Mira / @_mira___mira_ : There's obvious-seeming latency improvements I was wondering why nobody was doing: * Multiple tokens at once * Ability to exit early from the transformer stack * Use a cheaper computation to get “top 10 most likely” and speculatively compute all of them while the larger runs @teortaxestex : On the futility of small scaling experiments; or why frontier labs will out-innovate you Devastating figure [image] Aran Komatsuzaki / @arankomatsuzaki : Meta presents Better & Faster Large Language Models via Multi-token Prediction - training language models to predict multiple future tokens at once results in higher sample efficiency - up to 3x faster at inference https://arxiv.org/... [image] Yam Peleg / @yampeleg : Tried this already, it didn't improve the training for me. Probably is more efficient above certain amount of tokens that I didn't reach. Cool idea and very easy to try! Joey / @shxf0072 : tldr: > make n<=4 prediction lm head, > better perf & faster speculative decoding > with no overhead [image] @teortaxestex : Good question re Multi-Token Prediction. My answer: «at this model scale, 4 is optimal». And maybe at this data scale. It probably stacks well with curricula. Just as we're increasing seqlen throughout training, we may gradually increase text complexity *and* lookahead window. [image] Cody Blakeney / @code_star : I'm really excited about ideas like this, but before people get too worked up you should know this seems to be a domain specific intervention. Thats *ok* though. This might be a very useful piece of making code models or adapting models to be code models. I'm also willing to... [image] Tianle Cai / @tianle_cai : Wow, Medusa can be used for pre-training and leads to a better and faster generation! 😍 @teortaxestex : A detail I overlooked in Multi-Token Prediction: they test induction capability emergence using a synthetic dataset of children stories and small models, showing positive results until 30M. Wonder what dataset they mean, hm. Guess I have some intuition. [image] Archie / @archiexzzz : good read: better & faster llms via multi-token prediction by meta. training language models to predict multiple future tokens at once. https://arxiv.org/... [image] Saurabh Kumar / @drummatick : @archiexzzz Skipgram architecture did something similar with word2vec I haven't followed the development beyond that, but my guess is beyond n=5 the performance might be dropping and gains might not be uniform even for n>2 Kaizhao Liang / @kyleliang5 : old but interesting idea, this is the first time I see it working in the wild. https://arxiv.org/... what matters is the “intention”. If LLM wants to simulate human thought process, it needs t o be able to “know” what it wants to generate before it generates it. A mental draft... Ethan Mollick / @emollick : This may end up being a big deal: Usually LLMs just predict the next token in a sequence, one at a time, but if you have them predict the next several tokens at once you get significantly better performance, faster, and with no added costs. The gains are better for bigger models [image] LinkedIn: Jack FitzGerald : Instead of only training with next token prediction, some model parameters are instead distributed to n 1-layer heads on top of a shared trunk of n-1 layers … Shang-You T. : New paper “Better & Faster Large Language Models via Multi-token Prediction” can change direction of LLMs. … Forums: Hacker News : Better and Faster Large Language Models via Multi-Token Prediction
Context & Ripple Effects
Meta’s work arrives alongside its broader Llama push: it had detailed Llama 3’s 8B- and 70B-parameter releases weeks earlier, while this research targets the training objective rather than simply scaling model size. The proposed design keeps a shared trunk and adds independent output heads for several future tokens.
The idea moved beyond a paper shortly afterward, when Meta released research models using multi-token prediction. That progression makes the study a useful signal about an alternative route to improving model quality and training efficiency.
First-order effects
- Meta’s researchers gain a training recipe that, according to the study, improves performance and sample efficiency by supervising several future tokens at each position.
- Model builders experimenting with the method must alter the output layer and training loss to support multiple independent prediction heads, while retaining a shared model trunk.
Second-order effects
- If the reported gains reproduce across model families, teams face a trade-off between conventional next-token training and more complex objectives that may extract more learning from the same training data.
- The later release of pre-trained multi-token models gives researchers a concrete artifact for testing whether the paper’s training gains transfer to downstream use cases.
Third-order effects
- The work points toward model training becoming a more important optimization surface alongside parameter count: architectural and objective-level changes could matter as much as additional scale.
- Whether multi-token objectives become standard will depend on whether their training benefits persist under varied decoding and deployment conditions, rather than on benchmark gains alone.
The trend: LLM developers are increasingly seeking efficiency and quality gains through training objectives and architectures, not only larger models and more data.