/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

Meta claims both Llama 3 models beat similarly sized models like Gemini, Mistral, and Claude 3 on certain benchmarks; humans marked Llama 3 higher than GPT-3.5

The Verge Emilia David

Context & Ripple Effects

Llama 3 arrives amid a tightly contested model-performance race: several rival systems had already approached or exceeded GPT-4 on benchmarks, making relative results at comparable model sizes a key competitive claim rather than a standalone measure of capability.

Meta had previously positioned its Llama line around specialized performance, including long-context results for Llama 2 Long. The new claims broaden that positioning to general benchmark and human-preference comparisons against major proprietary and open-model rivals.

First-order effects

  • Meta gains marketing support for Llama 3 among developers weighing Gemini, Mistral, Claude 3, and GPT-3.5, though the reported advantage is limited to selected benchmarks and Meta's own evaluation framing.
  • The comparison puts pressure on rival model vendors to clarify performance at equivalent sizes and to compete on evaluation quality, cost, reliability, or deployment terms rather than headline model rankings alone.

Second-order effects

  • For model buyers, closer performance claims across vendors increase negotiating leverage and make workload-specific testing more important than choosing solely by a provider's flagship reputation.
  • Benchmark and human-preference results become more consequential inputs to developer adoption, while their methodology becomes a competitive issue—as later scrutiny of a non-public Llama 4 variant on a leaderboard illustrates.

Third-order effects

  • If capable models continue to converge at similar sizes, differentiation is likely to shift toward distribution, tooling, inference economics, and fit for particular workloads rather than a single aggregate benchmark lead.
  • The growing weight placed on benchmark claims may strengthen demand for more transparent, reproducible evaluations; rankings alone will remain an imperfect proxy for production performance.

The trend: This is one data point in the shift from a small set of clear model leaders toward a more competitive market where comparable capability increases buyer choice and raises the value of distribution and evaluation credibility.

Discussion

  • @zuck Mark Zuckerberg on threads
    To give a sense of performance, this 8B model is nearly as good as the biggest Llama 2 model.  This 70B model is around 82 MMLU with leading reasoning and math benchmarks.  The 400B+ model is currently around 85 MMLU but it's still training, so we expect it to lead on several ben…
  • @vishvanands Vishvanand Subramanian on threads
    Given Llama-3 400B is on par with the current best model Claude-3 Opus and it's still training, we can soon expect to see the dream of open source realized: The best model in the world is now free and open source.
  • @_kuanhoong_ @_kuanhoong_ on threads
    Llama 3 by Meta is here!  8B and 70B pretrained and instruction-tuned models are available. https://llama.meta.com/llama3/ Below is performance comparison Lllama-3 with Gemma, Gemini Pro 1.5, Mistal and Claude 3
  • @aiatmeta @aiatmeta on x
    Llama 3 delivers a major leap over Llama 2 and demonstrates SOTA performance on a wide range of industry benchmarks. The models also achieve substantially reduced false refusal rates, improved alignment and increased diversity in model responses — in addition to improved... [imag…
  • @ylecun Yann LeCun on x
    🥁 Llama3 is out 🥁 8B and 70B models available today. 8k context length.  Trained with 15 trillion tokens on a custom-built 24k GPU cluster.  Great performance on various benchmarks, with Llam3-8B doing better than Llama2-70B in some cases.  More versions are coming over the next …
  • @emollick Ethan Mollick on x
    Meta released their open source AI, Llama 3, today. As a key leader in LLMs, their models are often the most advanced open source ones out there. Based on benchmarks, the current model is not quite GPT-4 class, but their larger one (still training) will reach GPT-4 level. [image]
  • @mattshumer_ Matt Shumer on x
    The craziest LLaMA 3 reveal: The 400B+ version of the model is **on par with Claude 3 Opus**, and it's still training. Soon, we'll have a better-than-Opus, fully open-source model. The implications are huge. [image]
  • @mattshumer_ Matt Shumer on x
    Holy shit. LLaMA 3 70B cleanly beats Claude 3 Sonnet. Small enough to host at scale without breaking the bank. [image]