Meta releases Large Language Model Meta AI, or LLaMA, a foundational LLM designed to help AI researchers, available in sizes ranging from 7B to 65B parameters
As part of Meta's commitment to open science, today we are publicly releasing LLaMA (Large Language Model Meta AI) …
Meta AI
Context & Ripple Effects
LLaMA is the starting point for Meta’s model-release line: the related coverage later shows Meta broadening access with Llama 2’s research and commercial release and extending the family into coding through Code Llama. The initial range of model sizes matters because it gives researchers a common Meta foundation to build on rather than a single fixed model.
The subsequent arc also shows the strategy scaling upward: Meta later positioned Llama 3.1 405B as a frontier-level open-source model. That makes this release an early move in an increasingly consequential open-model distribution strategy, not an isolated research announcement.
First-order effects
- AI researchers gain access to Meta’s foundational model family across 7B-to-65B sizes, creating a shared base for research and downstream variants.
- Meta establishes LLaMA as a named model line that it can extend into broader-use and specialized releases.
Second-order effects
- The availability of LLaMA variants adds pressure to rival model providers to answer the expanding open-model ecosystem; related coverage reports OpenAI preparing an open-source LLM amid that proliferation.
- Developers seeking task-specific models gain a path toward Meta’s later family extensions, including Code Llama for generation and debugging and LLM Compiler for code optimization.
Third-order effects
- If Meta continues pairing broad model releases with larger and specialized variants, competition shifts from a single flagship model toward ecosystems built around reusable model families and downstream adaptation.
- The later move from LLaMA to Llama 3.1’s 405B open-source model suggests that open availability is becoming a competitive route for distributing increasingly capable foundation models.
The trend: Foundation-model competition is expanding from closed flagship systems toward open model families that researchers and developers can adapt across general and specialized tasks.
Related: AI Commons as Critical Infrastructure · Large Language Model Meta AI · Meta releases Llama 2 · Meta debuts Llama 3.1 405B · OpenAI prepares an open-source LLM
Related Coverage
- LLaMA: Open and Efficient Foundation Language Models Meta Research
- Meta rolls out AI language model LLaMA Reuters · Yuvraj Malik
- Meta releases Large Language Model Meta AI, or LLaMA, a foundational LLM designed to help AI researchers, available in sizes ranging from 7B to 65B parameters Bloomberg · Sarah Frier
- Meta has a new machine learning language model to remind you it does AI too The Verge
- Mark Zuckerberg announces Meta's new large language model as A.I. race heats up CNBC · Kif Leswing
- Meta unveils a new large language model that can run on a single GPU Ars Technica · Benj Edwards
- Today we're releasing a new state-of-the-art AI large language model called LLaMA designed to help researchers advance their work. … Mark Zuckerberg
- AI Daily Briefing 02/24 AI Daily Briefing
- Cause for a LLaMA? Meta reckons its smaller text-emitting AI is better than rivals The Register · Katyanna Quach
- Mark Zuckerberg Announces Meta's New Language Model Amidst ChatGPT Success Watcher Guru · Joshua Ramos
- Meta Announces Its Creating A New AI Large Language Model Called LLaMA Benzinga · Happy Mohamed
- Meta unveils new machine learning language model in ongoing AI investment Seeking Alpha · Jason Aycock
- Meta Offers Resources to Researchers Working in AI PYMNTS.com
- Meta announce LLaMA: A Foundational, 65-Billion-Parameter Large Language Model for researchers BigTechWire · Surur
- Zuckerberg Introduces Meta's Answer to ChatGPT, LLaMA Gizmodo · Lauren Leffer
- Facebook Parent Meta Wants to Show It's Still a Big Contender in AI Race CNET · Queenie Wong
- Meta's little LLaMA model comes with big benefits for AI researchers ZDNet · Stephanie Condon
- Mark Zuckerberg Announces Meta's New Large-Language Model AI ‘LLaMA’ SlashGear · Matt Salter
- Meta Platforms Releases New AI-Powered Language Model Appuals.com · Farhan Ali
- Mark Zuckerberg responds to the ChatGPT A.I. race with a new offering from Meta: Meet LLaMA Fortune
- Meta Debuts AI Language Model, But It's Only for Researchers PCMag · Michael Kan
- Meta reveals LLaMA language model as AI wars heat up BGR · Jacob Siegal
- Mark Zuckerberg just announced a new AI model ‘LLaMA,’ designed to help researchers make chatbots less ‘toxic’ Insider · Sindhu Sundar
- Meta to launch AI language model LLaMA to help researchers and take on ChatGPT Tech Startups · Nickie Louise
- Meta Releases LLaMA: A State-of-the-Art Foundational Language Model for AI Research Metaverse Post · Agne Cimermanaite
- Meta releases LLaMA to democratize access to large language AI models SiliconANGLE · Mike Wheatley
- Mark Zuckerberg says Meta is releasing LLaMA AI language model for researchers Neowin · Paul Hill
- Meta Joins AI Race With AI Language Model LLaMA Ethereum World News · Aditya Anand
Discussion
-
@guillaumelample
Guillaume Lample
on x
Today we release LLaMA, 4 foundation models ranging from 7B to 65B parameters. LLaMA-13B outperforms OPT and GPT-3 175B on most benchmarks. LLaMA-65B is competitive with Chinchilla 70B and PaLM 540B. The weights for all models are open and available at https://research.facebook.c…
-
@natfriedman
Nat Friedman
on x
It's interesting that Meta is releasing cool models but I'm not sure what they expect people to do with them given the license. https://twitter.com/... https://twitter.com/...
-
@schrep
Mike Schroepfer
on x
2 important details: a) Shows that smaller models trained with more data can outperform larger models (e.g. 13B outperforms GPT-3 175B) 2) The “larger” 65B model is competitive with best models - and is freely available to the research community! https://research.facebook.com/ ..…
-
@deliprao
Delip Rao
on x
FAIR releases LLaMA a new set of language models ranging from 7B to 65B that beats GPT3 and measures up to PaLM-540 in many tasks. Weights are available for research! https://research.facebook.com/ ... https://twitter.com/...
-
@metaai
@metaai
on x
Today we're publicly releasing LLaMA, a state-of-the-art foundational LLM, as part of our ongoing commitment to open science, transparency and democratized access to new research. Learn more & request access ➡️ https://ai.facebook.com/... https://twitter.com/...
-
@balajis
Balaji
on x
Are we decentralizing AI? Publicly available models. Trained on data with less lawsuit risk. Claims outperformance of GPT-3. Downloadable today... https://twitter.com/...
-
@ylecun
Yann LeCun
on x
LLaMA is a new *open-source*, high-performance large language model from Meta AI - FAIR. Meta is committed to open research and releases all the models the research community under a GPL v3 license. - Paper: https://research.facebook.com/ ... - Github: https://github.com/...
-
@ylecun
Yann LeCun
on x
@ESYudkowsky Your guess is indeed blind and hence wrong. Particularly on the whole “Meta has no people competent to do the tricky stuff...” Regardless, the difference between LLaMA and ChatGPT is fine tuning through human feedback It's expensive & time consuming, but not particul…
-
@jmdagdelen
John Dagdelen
on x
I commend FAIR for making their work truly accessible to other researchers. @ylecun and others are really speaking with their actions about the importance of maintaining open research. This is how to advance ML. https://twitter.com/...
-
@antirez
@antirez
on x
Sad to see the abuse of the term “open source” for something that can't be used for commercial uses and is sent only on request and after approval. https://twitter.com/...
-
@mkbhd
Marques Brownlee
on x
Google's large language model: “LaMDA” Large Model for Dialogue Applications Meta's new large language model: “LLaMA” [Screenshot of Mark Zuckerberg's announcement, and an image of a llama]
-
@kaushikpatnaik
@kaushikpatnaik
on x
Thoughts on the paper: 1) With LLaMA (+ also Galatica) it is now clear that we undertrain LLMs. More data and training longer than 1 epoch leads to improved performance on a variety of tasks https://twitter.com/...
-
@ylecun
Yann LeCun
on x
Generated by LLaMA. (prompt in bold). See appendix of the LLaMA paper: https://research.facebook.com/ ... https://twitter.com/...
-
@osanseviero
@osanseviero
on x
@GuillaumeLample Congratulations! Would be nice to have the models at @huggingface for increased discoverability and simpler usage. We have mechanisms for gated releases too!
-
@dystopiabreaker
@dystopiabreaker
on x
this is a big deal. it means gpt3 level performance in consumer gaming gpus https://twitter.com/...
-
@kaushikpatnaik
@kaushikpatnaik
on x
2) Given our knowledge of publicly available data for training models, it is increasingly obvious where the bottleneck is: data https://www.lesswrong.com/...
-
@kevinafischer
Kevin Fischer
on x
Facebook released a bunch of cool language models today. But why release them if no commercial use? A huge amount of interesting work using LLMs is being done in the commercial sector right now. They're enormously limiting the impact of their work by doing so. https://twitter.com…
-
@tri_dao
Tri Dao
on x
Amazing work by the team at Meta on training these foundation models! Glad to see FlashAttention being used to speed up the training process 🚀 https://twitter.com/...
-
@fluke_ellington
Teven Le Scao
on x
Meta proving once again that they're the best actor in this field out of all big tech, thanks for releasing all this hard work! https://twitter.com/...
-
@jffwng
Jeff Wang
on x
We are releasing LLaMA, a collection of foundational LLMs 7B to 65B in size. LLaMA-13B outperforms GPT-3 175B on most benchmarks. LLaMA-65B is competitive with the best models, Chinchilla70B and PaLM-540B Paper: https://research.facebook.com/ ... Apply for access: https://docs.go…
-
@esyudkowsky
Eliezer Yudkowsky
on x
I blindly guess, could be wrong, that this model will turn out sufficiently unimpressive in practice that nobody uses it for much. Basically based on a guess that more than benchmarks matter, and Meta has no people competent to do the tricky stuff needed to stay on current edge. …
-
@alexgraveley
Alex Graveley
on x
Open data only, great performance at fewer params. Practically the training data quality regime is the only secret sauce remaining in LLMs? https://twitter.com/...
-
@soumithchintala
Soumith Chintala
on x
LLaMa LLMs: small yet fierce from @GuillaumeLample and team at @MetaAI These are unaligned LLMs (so wont be ChatGPT quality). Compared to other base LLMs, they outperform relative to their size — LLama-65B is competitive with PaLM-540B. Read more at: https://research.facebook.com…
-
@antoinebordes
Antoine Bordes
on x
This release is great step towards understanding and reproducibility of LLMs! https://twitter.com/...
-
@debarghya_das
Deedy
on x
Facebook just dropped a 65B param LLM that outperforms GPT-3 and Google PaLM with ~12% the size of the latter by training on a high volume (1.4T tokens) of high quality text. More expensive to train, but much cheaper at inference! https://twitter.com/...
-
@francoisfleuret
@francoisfleuret
on x
Nice. FAIR is by far the most open AI lab among “the big corporations”. https://twitter.com/...
-
@armandjoulin
Armand Joulin
on x
Super excited to share new open LLMs from FAIR with our research community. Particularly, the LLaMA-13B is competitive with GPT-3, despite being 10x smaller. https://twitter.com/...
-
@wintonark
Brett Winton
on x
Meta open sources a language model built on publicly available data that outperforms GPT-3 (at less than 1/10th the inference cost) ! https://twitter.com/...
-
@wintonark
Brett Winton
on x
Having a data advantage in AI could result in a massive performance/cost advantage. Relative to 800 billion tokens, 1.2 trillion tokens allows you to reduce training cost by 1/3rd and reduce inference costs by half for the same performance output. https://twitter.com/... https://…
-
@arankomatsuzaki
Aran Komatsuzaki
on x
The link to the repo is here: https://github.com/... Fill in the form to download the checkpoints. The 65B model performs almost on par with PaLM 540B on many tasks thanks to its Chinchilla-optimal training and various tricks. https://twitter.com/...
-
@simonw
Simon Willison
on x
From the paper: “For instance, LLaMA-13B outperforms GPT-3 on most bench- marks, despite being 10× smaller. We believe that this model will help democratize the access and study of LLMs, since it can be run on a single GPU.” https://twitter.com/...
-
@naveengrao
Naveen Rao
on x
This is very cool! Thanks @MetaAI! From what I can tell, this is a GNU license, which means fine-tuned models all must be upstreamed. For corporates, this is challenging as fine-tuned models on proprietary data will need to stay with the company. @MetaAI any idea if there are... …
-
@abacaj
Anton
on x
Meta would have a much bigger impact if we could use it for non research purposes and if it wasn't gated to researchers https://twitter.com/...
-
@yabhishekhd
Abhishek Yadav
on x
Meta launches AI large language model called LLaMA designed to help researchers advance their work. LLMs have shown a lot of promise in generating text, having conversations, summarizing written material, and more complicated tasks like solving math theorems or (1/2)
-
@schrep
Mike Schroepfer
on x
If you are paying attention to AI you should follow this - the core advance powering many AI powered experiences is a foundational model. These are very difficult and expensive to train - so access is limited. This brings state of the art models to the entire research community h…