/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

Meta releases Large Language Model Meta AI, or LLaMA, a foundational LLM designed to help AI researchers, available in sizes ranging from 7B to 65B parameters

As part of Meta's commitment to open science, today we are publicly releasing LLaMA (Large Language Model Meta AI) …

Meta AI

Context & Ripple Effects

LLaMA is the starting point for Meta’s model-release line: the related coverage later shows Meta broadening access with Llama 2’s research and commercial release and extending the family into coding through Code Llama. The initial range of model sizes matters because it gives researchers a common Meta foundation to build on rather than a single fixed model.

The subsequent arc also shows the strategy scaling upward: Meta later positioned Llama 3.1 405B as a frontier-level open-source model. That makes this release an early move in an increasingly consequential open-model distribution strategy, not an isolated research announcement.

First-order effects

  • AI researchers gain access to Meta’s foundational model family across 7B-to-65B sizes, creating a shared base for research and downstream variants.
  • Meta establishes LLaMA as a named model line that it can extend into broader-use and specialized releases.

Second-order effects

  • The availability of LLaMA variants adds pressure to rival model providers to answer the expanding open-model ecosystem; related coverage reports OpenAI preparing an open-source LLM amid that proliferation.
  • Developers seeking task-specific models gain a path toward Meta’s later family extensions, including Code Llama for generation and debugging and LLM Compiler for code optimization.

Third-order effects

  • If Meta continues pairing broad model releases with larger and specialized variants, competition shifts from a single flagship model toward ecosystems built around reusable model families and downstream adaptation.
  • The later move from LLaMA to Llama 3.1’s 405B open-source model suggests that open availability is becoming a competitive route for distributing increasingly capable foundation models.

The trend: Foundation-model competition is expanding from closed flagship systems toward open model families that researchers and developers can adapt across general and specialized tasks.

Discussion

  • @guillaumelample Guillaume Lample on x
    Today we release LLaMA, 4 foundation models ranging from 7B to 65B parameters. LLaMA-13B outperforms OPT and GPT-3 175B on most benchmarks. LLaMA-65B is competitive with Chinchilla 70B and PaLM 540B. The weights for all models are open and available at https://research.facebook.c…
  • @natfriedman Nat Friedman on x
    It's interesting that Meta is releasing cool models but I'm not sure what they expect people to do with them given the license. https://twitter.com/... https://twitter.com/...
  • @schrep Mike Schroepfer on x
    2 important details: a) Shows that smaller models trained with more data can outperform larger models (e.g. 13B outperforms GPT-3 175B) 2) The “larger” 65B model is competitive with best models - and is freely available to the research community! https://research.facebook.com/ ..…
  • @deliprao Delip Rao on x
    FAIR releases LLaMA a new set of language models ranging from 7B to 65B that beats GPT3 and measures up to PaLM-540 in many tasks. Weights are available for research! https://research.facebook.com/ ... https://twitter.com/...
  • @metaai @metaai on x
    Today we're publicly releasing LLaMA, a state-of-the-art foundational LLM, as part of our ongoing commitment to open science, transparency and democratized access to new research. Learn more & request access ➡️ https://ai.facebook.com/... https://twitter.com/...
  • @balajis Balaji on x
    Are we decentralizing AI? Publicly available models. Trained on data with less lawsuit risk. Claims outperformance of GPT-3. Downloadable today... https://twitter.com/...
  • @ylecun Yann LeCun on x
    LLaMA is a new *open-source*, high-performance large language model from Meta AI - FAIR. Meta is committed to open research and releases all the models the research community under a GPL v3 license. - Paper: https://research.facebook.com/ ... - Github: https://github.com/...
  • @ylecun Yann LeCun on x
    @ESYudkowsky Your guess is indeed blind and hence wrong. Particularly on the whole “Meta has no people competent to do the tricky stuff...” Regardless, the difference between LLaMA and ChatGPT is fine tuning through human feedback It's expensive & time consuming, but not particul…
  • @jmdagdelen John Dagdelen on x
    I commend FAIR for making their work truly accessible to other researchers. @ylecun and others are really speaking with their actions about the importance of maintaining open research. This is how to advance ML. https://twitter.com/...
  • @antirez @antirez on x
    Sad to see the abuse of the term “open source” for something that can't be used for commercial uses and is sent only on request and after approval. https://twitter.com/...
  • @mkbhd Marques Brownlee on x
    Google's large language model: “LaMDA” Large Model for Dialogue Applications Meta's new large language model: “LLaMA” [Screenshot of Mark Zuckerberg's announcement, and an image of a llama]
  • @kaushikpatnaik @kaushikpatnaik on x
    Thoughts on the paper: 1) With LLaMA (+ also Galatica) it is now clear that we undertrain LLMs. More data and training longer than 1 epoch leads to improved performance on a variety of tasks https://twitter.com/...
  • @ylecun Yann LeCun on x
    Generated by LLaMA. (prompt in bold). See appendix of the LLaMA paper: https://research.facebook.com/ ... https://twitter.com/...
  • @osanseviero @osanseviero on x
    @GuillaumeLample Congratulations! Would be nice to have the models at @huggingface for increased discoverability and simpler usage. We have mechanisms for gated releases too!
  • @dystopiabreaker @dystopiabreaker on x
    this is a big deal. it means gpt3 level performance in consumer gaming gpus https://twitter.com/...
  • @kaushikpatnaik @kaushikpatnaik on x
    2) Given our knowledge of publicly available data for training models, it is increasingly obvious where the bottleneck is: data https://www.lesswrong.com/...
  • @kevinafischer Kevin Fischer on x
    Facebook released a bunch of cool language models today. But why release them if no commercial use? A huge amount of interesting work using LLMs is being done in the commercial sector right now. They're enormously limiting the impact of their work by doing so. https://twitter.com…
  • @tri_dao Tri Dao on x
    Amazing work by the team at Meta on training these foundation models! Glad to see FlashAttention being used to speed up the training process 🚀 https://twitter.com/...
  • @fluke_ellington Teven Le Scao on x
    Meta proving once again that they're the best actor in this field out of all big tech, thanks for releasing all this hard work! https://twitter.com/...
  • @jffwng Jeff Wang on x
    We are releasing LLaMA, a collection of foundational LLMs 7B to 65B in size. LLaMA-13B outperforms GPT-3 175B on most benchmarks. LLaMA-65B is competitive with the best models, Chinchilla70B and PaLM-540B Paper: https://research.facebook.com/ ... Apply for access: https://docs.go…
  • @esyudkowsky Eliezer Yudkowsky on x
    I blindly guess, could be wrong, that this model will turn out sufficiently unimpressive in practice that nobody uses it for much. Basically based on a guess that more than benchmarks matter, and Meta has no people competent to do the tricky stuff needed to stay on current edge. …
  • @alexgraveley Alex Graveley on x
    Open data only, great performance at fewer params. Practically the training data quality regime is the only secret sauce remaining in LLMs? https://twitter.com/...
  • @soumithchintala Soumith Chintala on x
    LLaMa LLMs: small yet fierce from @GuillaumeLample and team at @MetaAI These are unaligned LLMs (so wont be ChatGPT quality). Compared to other base LLMs, they outperform relative to their size — LLama-65B is competitive with PaLM-540B. Read more at: https://research.facebook.com…
  • @antoinebordes Antoine Bordes on x
    This release is great step towards understanding and reproducibility of LLMs! https://twitter.com/...
  • @debarghya_das Deedy on x
    Facebook just dropped a 65B param LLM that outperforms GPT-3 and Google PaLM with ~12% the size of the latter by training on a high volume (1.4T tokens) of high quality text. More expensive to train, but much cheaper at inference! https://twitter.com/...
  • @francoisfleuret @francoisfleuret on x
    Nice. FAIR is by far the most open AI lab among “the big corporations”. https://twitter.com/...
  • @armandjoulin Armand Joulin on x
    Super excited to share new open LLMs from FAIR with our research community. Particularly, the LLaMA-13B is competitive with GPT-3, despite being 10x smaller. https://twitter.com/...
  • @wintonark Brett Winton on x
    Meta open sources a language model built on publicly available data that outperforms GPT-3 (at less than 1/10th the inference cost) ! https://twitter.com/...
  • @wintonark Brett Winton on x
    Having a data advantage in AI could result in a massive performance/cost advantage. Relative to 800 billion tokens, 1.2 trillion tokens allows you to reduce training cost by 1/3rd and reduce inference costs by half for the same performance output. https://twitter.com/... https://…
  • @arankomatsuzaki Aran Komatsuzaki on x
    The link to the repo is here: https://github.com/... Fill in the form to download the checkpoints. The 65B model performs almost on par with PaLM 540B on many tasks thanks to its Chinchilla-optimal training and various tricks. https://twitter.com/...
  • @simonw Simon Willison on x
    From the paper: “For instance, LLaMA-13B outperforms GPT-3 on most bench- marks, despite being 10× smaller. We believe that this model will help democratize the access and study of LLMs, since it can be run on a single GPU.” https://twitter.com/...
  • @naveengrao Naveen Rao on x
    This is very cool! Thanks @MetaAI! From what I can tell, this is a GNU license, which means fine-tuned models all must be upstreamed. For corporates, this is challenging as fine-tuned models on proprietary data will need to stay with the company. @MetaAI any idea if there are... …
  • @abacaj Anton on x
    Meta would have a much bigger impact if we could use it for non research purposes and if it wasn't gated to researchers https://twitter.com/...
  • @yabhishekhd Abhishek Yadav on x
    Meta launches AI large language model called LLaMA designed to help researchers advance their work. LLMs have shown a lot of promise in generating text, having conversations, summarizing written material, and more complicated tasks like solving math theorems or (1/2)
  • @schrep Mike Schroepfer on x
    If you are paying attention to AI you should follow this - the core advance powering many AI powered experiences is a foundational model. These are very difficult and expensive to train - so access is limited. This brings state of the art models to the entire research community h…