xAI says Grok-3 outperforms Gemini-2 Pro, DeepSeek-V3, Claude 3.5 Sonnet, and GPT-4o in some benchmarks; Musk says xAI's mission is to “understand the universe”
by voting with their feet, and the markets— by voting to give Grok dollars It's clearly a SOTA model, and folks who threw shade at Grok research team were mistaken [image] Aaron Levie / @levie : Grok 3 benchmarks show the jump in AI capability you get when it spends more time on a task. In the future, you will be able to solve any given problem in the world just by throwing more compute at it. [video]
Grok-3 reasoning is not released yet The version that is released wasn't doing so well on their three self-reported benchmarks Technically there is nothing to evaluate or test yet! So will just have to wait 🤷♀️
Another thing Grok 3 highlights is the urgent need for better batteries of tests and independent testing authorities. Public benchmarks are both “meh” and saturated, leaving a lot of AI testing to be like food reviews, based on taste. If AI is critical to to work, we need more.
Is ~log(15)x improvement in these benchmarks worth it for ~15x cluster scaling? In the end users will decide— by voting with their feet, and the markets— by voting to give Grok dollars It's clearly a SOTA model, and folks who threw shade at Grok research team were mistaken [image…
Grok 3 benchmarks show the jump in AI capability you get when it spends more time on a task. In the future, you will be able to solve any given problem in the world just by throwing more compute at it. [video]