/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

xAI unveils Grok 4.1, saying its hallucination rate is 3x less compared to its previous models and Grok 4.1 Thinking holds the top spot on LMArena's Text Arena

Grok 4.1 is now available to all users on grok.com, 𝕏, and the iOS and Android apps.  It is rolling out immediately …

xAI

Context & Ripple Effects

xAI has moved quickly from the Grok-3 reasoning-model beta to Grok 4, which was introduced with multimodal and coding features. The new release makes reliability—not only capability breadth—the central claimed improvement.

The update follows xAI’s push into premium performance tiers through Grok 4 Heavy and its high-priced subscription plan, while Grok 4’s launch was also associated with a sharp short-term increase in iOS revenue. Broad availability across xAI’s existing surfaces gives the company a direct channel to test whether quality gains sustain usage.

First-order effects

  • Users of grok.com, X, and Grok’s mobile apps receive Grok 4.1 immediately, with xAI positioning its lower hallucination-rate claim as a practical quality upgrade over prior Grok models.
  • Grok 4.1 Thinking’s Text Arena lead gives xAI a current third-party leaderboard credential alongside its own reliability claim, strengthening its product messaging.

Second-order effects

  • Competing model providers face added pressure to demonstrate both benchmark performance and factual reliability, rather than relying on one measure alone.
  • For xAI, the release tests whether improvements that follow the Grok 4-driven iOS revenue jump can translate into more durable engagement or subscription appeal; the available coverage does not establish that outcome.

Third-order effects

  • If model updates increasingly compete on error reduction, evaluation will shift toward the cost per dependable answer or task, not just peak benchmark placement.
  • Distribution through a social platform, web service, and native apps makes model quality upgrades easier to place before existing users, reinforcing an AI distribution advantage—but independent validation of reliability claims will remain important.

The trend: Frontier AI competition is shifting from headline benchmark wins toward frequent, widely distributed upgrades that aim to make model outputs more dependable in everyday use.

Discussion

  • @arena @arena on x
    🚨Text Leaderboard Update @xAI's Grok 4.1 (thinking) and Grok 4.1 have scaled new heights in the most competitive Text Arena: 🔹Grok 4.1 (thinking) lands at #1 with a score of 1483 🔹Grok 4.1 follows at #2 with a score of 1465 On the Arena Expert leaderboard: 🔸Grok 4.1 [image]
  • @minimaxir Max Woolf on x
    I can confirm Grok 4.1 has effectively no content filters: even on the web UI which should have its own safety prompts, it's *extremely* permissive and I suspect that the other safety filters in its model card can be defeated. Also, wtf at those next prompt suggestions. [image]
  • @chrisparkx Chris Park on x
    Grok 4.1 is an insanely good (the best thus far) model sir. Congrats and great work by core xAI team in shipping this! 🚀🚀🚀
  • @thekitze @thekitze on x
    “we have gemini 3 at home”
  • @flavioad Flavio Adamo on x
    kinda wondering why the official Grok 4.1 blog post didn't include the SWE-bench score feels a bit suspicious ngl
  • @zeffmax Max Zeff on x
    The model card for Grok 4.1 is wild... xAI seems to say — extremely vaguely — that they improved on sycophancy, but it actually got more than 2x worse? [image]
  • @leothecurious @leothecurious on x
    i don't even notice anymore. i think this is it. we're well into the second half of the sigmoid for this breakthrough. fingers crossed for gemini 3 and i'm sure it'll be great but idk how much more we can expect to squeeze out of those things. i'm sure we'll keep seeing numbers
  • @mark_k Mark Kretschmann on x
    Grok 4.1 first impressions: * Very different writing style, much more personal. * Much improved image understanding, sees even small details. * Comes in two flavors, normal and “thinking”. * Quite fast, even the thinking version.
  • @arena @arena on x
    Though preliminary, @xAI's Grok 4.1 (thinking) also lands #1 on the Expert leaderboard and shines in the following overview categories: 🌶️Hard Prompts 💻Coding 📝Instruction Following ✍️Creative Writing [image]
  • @arena @arena on x
    On the Occupational Leaderboard, Grok 4.1 (thinking) shows top strength in nearly all fields. 💻 Software & IT Services ✍️ Writing, Literature, & Language 🔬 Life, Physical, & Social Science 🎭 Entertainment, Sports, & Media 📈 Business, Management, & Financial Ops ⚖️ Legal &
  • @miles_brundage Miles Brundage on x
    😬 https://x.com/... [image]
  • @yuhu_ai_ @yuhu_ai_ on x
    Grok4.1 thinking and non thinking sitting at #1 and #2, 31 elo scores above others.
  • @koltregaskes @koltregaskes on x
    Grok 4.1 delivers gains in creative writing, emotional intelligence and personality coherence while retaining raw intelligence. - Thinking mode #1 on LMArena at 1483* Elo, non-reasoning #2 at 1465* Elo and beats every other model's reasoning version - Tops EQ-Bench3 for [image]
  • @mark_k Mark Kretschmann on x
    IMO @xai should remove the non-thinking version of Grok 4.1 completely, @elonmusk. The thinking version is fast enough now and much better!
  • @ai_for_success AshutoshShrivastava on x
    First quick test of Grok 4.1 involving doing research and coming back with a proper detailed response. The questions I asked are things I already know, I just wanted to see the quality of its response and it was pretty good and detailed. I tested a few more points and it [video]
  • @scaling01 @scaling01 on x
    Grok 4.1 absolutely smashes all other models on lmarena with an Elo of 1483 it comes with higher emotional intelligence, better creative writing and less hallucinations [image]
  • @nearlydaniel Daniel on x
    The whole team cooked super hard on 4.1, really excited to share it with you all! Model personality and quality, low latency, reliability up and down the stack - lots of improvements. Try out Grok 4.1 and let me know what you think 🚀🚀
  • @yiwenyuan98 Yiwen Yuan on x
    Grok 4.1 can now do technical analysis on stocks very quickly and accurately with code execution. Good news to all the stock fans! https://grok.com/... [image]
  • @litianleli @litianleli on x
    Grok 4.1 is PEAK post-training. We unlocked a ton of new recipes and pushed the model to absolute frontier performance across a bunch of hard-to-verify, general domains: emotional intelligence, hallucination rate, chat, creative writing, latency, and efficiency. Making big
  • @adityagupta Aditya Gupta on x
    Over the past few weeks, we have been working on the post-training RL for sharpening model's alignment with users' preferences in conjunction with model's capabilities and intelligence. It has been an amazing learning journey about the recipe, product, user signals, style,
  • @elonmusk Elon Musk on x
    Grok 4.1 just released. You should notice a significant increase in speed and quality.
  • @xai @xai on x
    Introducing Grok 4.1, a frontier model that sets a new standard for conversational intelligence, emotional understanding, and real-world helpfulness. Grok 4.1 is available for free on https://grok.com/, https://grok.x.com/ and our mobile apps. https://x.ai/...
  • @billyuchenlin Bill Yuchen Lin on x
    Grok 4.1 is our latest model! Higher IQ & EQ, better at creative tasks, and less hallucination. 🚀 Download @grok to try it out! https://x.ai/...
  • @olaindeach.oddling.nl @olaindeach.oddling.nl on bluesky
    Thats probably 16x less likely than its owner...  [embedded post]
  • r/MyBoyfriendIsAI r on reddit
    Grok 4.1 |  xAI