xAI unveils Grok 4.1, saying its hallucination rate is 3x less compared to its previous models and Grok 4.1 Thinking holds the top spot on LMArena's Text Arena
Grok 4.1 is now available to all users on grok.com, 𝕏, and the iOS and Android apps. It is rolling out immediately …
xAI
Context & Ripple Effects
xAI has moved quickly from the Grok-3 reasoning-model beta to Grok 4, which was introduced with multimodal and coding features. The new release makes reliability—not only capability breadth—the central claimed improvement.
The update follows xAI’s push into premium performance tiers through Grok 4 Heavy and its high-priced subscription plan, while Grok 4’s launch was also associated with a sharp short-term increase in iOS revenue. Broad availability across xAI’s existing surfaces gives the company a direct channel to test whether quality gains sustain usage.
First-order effects
Users of grok.com, X, and Grok’s mobile apps receive Grok 4.1 immediately, with xAI positioning its lower hallucination-rate claim as a practical quality upgrade over prior Grok models.
Grok 4.1 Thinking’s Text Arena lead gives xAI a current third-party leaderboard credential alongside its own reliability claim, strengthening its product messaging.
Second-order effects
Competing model providers face added pressure to demonstrate both benchmark performance and factual reliability, rather than relying on one measure alone.
For xAI, the release tests whether improvements that follow the Grok 4-driven iOS revenue jump can translate into more durable engagement or subscription appeal; the available coverage does not establish that outcome.
Third-order effects
If model updates increasingly compete on error reduction, evaluation will shift toward the cost per dependable answer or task, not just peak benchmark placement.
Distribution through a social platform, web service, and native apps makes model quality upgrades easier to place before existing users, reinforcing an AI distribution advantage—but independent validation of reliability claims will remain important.
The trend: Frontier AI competition is shifting from headline benchmark wins toward frequent, widely distributed upgrades that aim to make model outputs more dependable in everyday use.
🚨Text Leaderboard Update @xAI's Grok 4.1 (thinking) and Grok 4.1 have scaled new heights in the most competitive Text Arena: 🔹Grok 4.1 (thinking) lands at #1 with a score of 1483 🔹Grok 4.1 follows at #2 with a score of 1465 On the Arena Expert leaderboard: 🔸Grok 4.1 [image]
I can confirm Grok 4.1 has effectively no content filters: even on the web UI which should have its own safety prompts, it's *extremely* permissive and I suspect that the other safety filters in its model card can be defeated. Also, wtf at those next prompt suggestions. [image]
The model card for Grok 4.1 is wild... xAI seems to say — extremely vaguely — that they improved on sycophancy, but it actually got more than 2x worse? [image]
i don't even notice anymore. i think this is it. we're well into the second half of the sigmoid for this breakthrough. fingers crossed for gemini 3 and i'm sure it'll be great but idk how much more we can expect to squeeze out of those things. i'm sure we'll keep seeing numbers
Grok 4.1 first impressions: * Very different writing style, much more personal. * Much improved image understanding, sees even small details. * Comes in two flavors, normal and “thinking”. * Quite fast, even the thinking version.
Though preliminary, @xAI's Grok 4.1 (thinking) also lands #1 on the Expert leaderboard and shines in the following overview categories: 🌶️Hard Prompts 💻Coding 📝Instruction Following ✍️Creative Writing [image]
On the Occupational Leaderboard, Grok 4.1 (thinking) shows top strength in nearly all fields. 💻 Software & IT Services ✍️ Writing, Literature, & Language 🔬 Life, Physical, & Social Science 🎭 Entertainment, Sports, & Media 📈 Business, Management, & Financial Ops ⚖️ Legal &
Grok 4.1 delivers gains in creative writing, emotional intelligence and personality coherence while retaining raw intelligence. - Thinking mode #1 on LMArena at 1483* Elo, non-reasoning #2 at 1465* Elo and beats every other model's reasoning version - Tops EQ-Bench3 for [image]
First quick test of Grok 4.1 involving doing research and coming back with a proper detailed response. The questions I asked are things I already know, I just wanted to see the quality of its response and it was pretty good and detailed. I tested a few more points and it [video]
Grok 4.1 absolutely smashes all other models on lmarena with an Elo of 1483 it comes with higher emotional intelligence, better creative writing and less hallucinations [image]
The whole team cooked super hard on 4.1, really excited to share it with you all! Model personality and quality, low latency, reliability up and down the stack - lots of improvements. Try out Grok 4.1 and let me know what you think 🚀🚀
Grok 4.1 can now do technical analysis on stocks very quickly and accurately with code execution. Good news to all the stock fans! https://grok.com/... [image]
Grok 4.1 is PEAK post-training. We unlocked a ton of new recipes and pushed the model to absolute frontier performance across a bunch of hard-to-verify, general domains: emotional intelligence, hallucination rate, chat, creative writing, latency, and efficiency. Making big
Over the past few weeks, we have been working on the post-training RL for sharpening model's alignment with users' preferences in conjunction with model's capabilities and intelligence. It has been an amazing learning journey about the recipe, product, user signals, style,
Introducing Grok 4.1, a frontier model that sets a new standard for conversational intelligence, emotional understanding, and real-world helpfulness. Grok 4.1 is available for free on https://grok.com/, https://grok.x.com/ and our mobile apps. https://x.ai/...