Meta claims both Llama 3 models beat similarly sized models like Gemini, Mistral, and Claude 3 on certain benchmarks; humans marked Llama 3 higher than GPT-3.5
Meta had previously positioned its Llama line around specialized performance, including long-context results for Llama 2 Long. The new claims broaden that positioning to general benchmark and human-preference comparisons against major proprietary and open-model rivals.
First-order effects
Meta gains marketing support for Llama 3 among developers weighing Gemini, Mistral, Claude 3, and GPT-3.5, though the reported advantage is limited to selected benchmarks and Meta's own evaluation framing.
The comparison puts pressure on rival model vendors to clarify performance at equivalent sizes and to compete on evaluation quality, cost, reliability, or deployment terms rather than headline model rankings alone.
Second-order effects
For model buyers, closer performance claims across vendors increase negotiating leverage and make workload-specific testing more important than choosing solely by a provider's flagship reputation.
Benchmark and human-preference results become more consequential inputs to developer adoption, while their methodology becomes a competitive issue—as later scrutiny of a non-public Llama 4 variant on a leaderboard illustrates.
Third-order effects
If capable models continue to converge at similar sizes, differentiation is likely to shift toward distribution, tooling, inference economics, and fit for particular workloads rather than a single aggregate benchmark lead.
The growing weight placed on benchmark claims may strengthen demand for more transparent, reproducible evaluations; rankings alone will remain an imperfect proxy for production performance.
The trend: This is one data point in the shift from a small set of clear model leaders toward a more competitive market where comparable capability increases buyer choice and raises the value of distribution and evaluation credibility.
To give a sense of performance, this 8B model is nearly as good as the biggest Llama 2 model. This 70B model is around 82 MMLU with leading reasoning and math benchmarks. The 400B+ model is currently around 85 MMLU but it's still training, so we expect it to lead on several ben…
Given Llama-3 400B is on par with the current best model Claude-3 Opus and it's still training, we can soon expect to see the dream of open source realized: The best model in the world is now free and open source.
Llama 3 by Meta is here! 8B and 70B pretrained and instruction-tuned models are available. https://llama.meta.com/llama3/ Below is performance comparison Lllama-3 with Gemma, Gemini Pro 1.5, Mistal and Claude 3
Llama 3 delivers a major leap over Llama 2 and demonstrates SOTA performance on a wide range of industry benchmarks. The models also achieve substantially reduced false refusal rates, improved alignment and increased diversity in model responses — in addition to improved... [imag…
🥁 Llama3 is out 🥁 8B and 70B models available today. 8k context length. Trained with 15 trillion tokens on a custom-built 24k GPU cluster. Great performance on various benchmarks, with Llam3-8B doing better than Llama2-70B in some cases. More versions are coming over the next …
Meta released their open source AI, Llama 3, today. As a key leader in LLMs, their models are often the most advanced open source ones out there. Based on benchmarks, the current model is not quite GPT-4 class, but their larger one (still training) will reach GPT-4 level. [image]
The craziest LLaMA 3 reveal: The 400B+ version of the model is **on par with Claude 3 Opus**, and it's still training. Soon, we'll have a better-than-Opus, fully open-source model. The implications are huge. [image]