/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

SpaceXAI releases Grok 4.6, saying it matches GPT-5.6 Sol on the Artificial Analysis Intelligence Index, and prices it at $2/1M input and $6/1M output tokens

Grok 4.6 builds on Grok 4.5 with a particular focus on long-running agents and more ambitious interactive and visual work.  —  Try for free

xAI

Context & Ripple Effects

Grok 4.5 had already reached Grok Build, Cursor and the SpaceXAI console at the same published token rates, though it was not available in the EU. Grok 4.6 therefore pairs a new workload focus with Grok 4.5's existing $2/$6 pricing.

The release extends a line that previously claimed benchmark leadership for Grok 4 and introduced Grok 4 Heavy's multi-agent variant. The emphasis has shifted toward sustained agent runs and interactive or visual work rather than a benchmark position alone.

First-order effects

Second-order effects

  • For teams selecting models for agent workloads, the comparison shifts from token pricing alone toward whether Grok 4.6's stated capabilities deliver better results on extended, interactive tasks.
  • SpaceXAI's use of the Artificial Analysis Intelligence Index makes that benchmark a more prominent reference point in its competition with GPT-5.6 Sol.

Third-order effects

  • If model vendors hold token rates steady across upgrades while targeting long-running agents, procurement will increasingly center on cost per completed task rather than cost per token.
  • Grok's progression from a benchmark-leading model to a multi-agent offering and now long-running-agent positioning points toward competition over agent execution, not only single-model scores.

The trend: Frontier-model competition is moving from headline benchmark claims toward agentic performance delivered at stable inference prices.

Discussion

  • @spacexai @spacexai on x
    Introducing Grok 4.6. It delivers frontier intelligence and is a significant improvement over Grok 4.5 at the same price. [image]
  • @artificialanlys @artificialanlys on x
    SpaceXAI's Grok 4.6 scores 61 on the Artificial Analysis Intelligence Index, joining the frontier in line with GPT-5.6 Sol, with standout agentic performance at lower cost...Key takeaways: ➤ Grok 4.6 joins the frontier of the Artificial Analysis Intelligence Index...Strong agenti…
  • @elonmusk Elon Musk on x
    Grok 4.6 is objectively #1 when considering intelligence, speed & cost [image]
  • @cognition @cognition on x
    Grok 4.6 is now available in Devin. Grok 4.6 marks a significant improvement over Grok 4.5, surpassing GPT-5.6 Sol, behind only Opus 5 and Fable 5. [image]
  • @elonmusk Elon Musk on x
    @cognition Grok 4.7 will exceed all current models. That said, Anthropic is a great company and will probably release improved models soon. However, the SpaceX training corpus is so awesome & unique that I would be shocked if any model is better at real-world engineering than 4.7…
  • @arena @arena on x
    Grok 4.6 is here!  This release by @SpaceXAI just landed in the Code Arena: WebDev at #7 with 1618 pts.  Grok 4.6 (High) is a big jump from Grok 4.5, which sits at #13 with 1553 pts.  It's now on par with GPT-5.6 Sol xHigh (1622 pts) and Claude Fable 5 (1627 pts).  All three curr…
  • @elonmusk Elon Musk on x
    Grok 4.6 is now out 🚀🚀🚀 Smart, fast & amazing bang for buck!
  • @patrickmoorhead Patrick Moorhead on x
    Great to see the competition. It'll take while to gain the trust that this isn't a one hit wonder before people disconnect from what they're using today.
  • @elonmusk Elon Musk on x
    Try Grok 4.6 on tough real-world tasks!
  • @shaunmmaguire Shaun Maguire on x
    A good day to repost this! And as a reminder... it's just the beginning Buckle up 🚀
  • @synthwavedd Leo on x
    xAI have made an incredible comeback From the days of Grok 4 -> 4.3, where they were trailing the frontier by far, they're now arguably the 3rd best lab in the world - behind only Anthropic and OpenAI And with this minor version jump yielding +5 on AA's Intelligence Index, and Gr…
  • @theskory Boris Skorobogaty on x
    Grok 4.6 is an excellent model. I've been using it heavily for the past couple of weeks and it handles everything from simple coding & code review all the way to designing and debugging complex systems with ease. What stands out even more is how strong it is outside of coding -
  • @iterintellectus Vittorio on x
    so they just... caught up? in three years?! grok is now frontier how does he do it?!
  • @gavinsbaker Gavin Baker on x
    Grok 4.6 is roughly the same performance as Fable 5 Max at an 85% discount. 80% cheaper for input tokens and 88% cheaper for output tokens. Pareto dominant. Grok 4.7 will be significantly better as is a much larger model with the Cursor and SpaceX data included in pretraining. [i…
  • @ivanzhouyq Ivan Zhou on x
    We worked with @SpaceXAI to evaluate Grok 4.6 on the latest OfficeQA Pro V2 from @DbrxMosaicAI. It achieves the SOTA performance with @databricks's Genie harness! The model is strongest on our document understanding and data reasoning tasks, and it is a very efficient driver! [im…
  • @levelsio @levelsio on x
    I switched https://ideasai.com/'s auto app and landing builder to @xai Grok now Grok 4.6 is apparently as good as Fable 5 (without the constant preaching and blocking you from doing anything due to SeCuRiTy) [image]
  • @afinetheorem Kevin A. Bryan on x
    On my vision+logic benchmark, Grok 4.6 (and Qwen3.8 Max) remain very mediocre - between Gemini 2.0 Flash Lite and GPT 5.6 Luna. Best still Gemini models - good at this type of task in many benchmarks - Fable & Muse 1.2. GPT models good but haven't gotten better since o3 > 1yr ago…
  • @_nathancalvin Nathan Calvin on x
    Is this all the safety information they plan to put out for this release? Remarkably far behind peers at Meta, Anthropic, OpenAI, and Google [image]
  • @tenobrus @tenobrus on x
    goddamn. never thought i'd see the day where xai really decisively flips deepmind
  • @davidlee @davidlee on x
    85% cheaper is shopping outlet prices Grok is also the fastest with the most relevant unique data for my work Biggest issue has been the UX-harness which @bot may have solved yesterday
  • @spacexai @spacexai on x
    Grok 4.6 is faster than comparable models and can handle much more challenging tasks than Grok 4.5. It's half the price of other frontier models at $2/M input and $6/M output tokens.
  • @elonmusk Elon Musk on x
    Grok 4.6 reaches #1 on @databricks
  • @elonmusk Elon Musk on x
    Grok 4.6 reaches 1753 ELO
  • @nicdunz Nic on x
    ok this is insane. at effectively the same overall intelligence score, grok 4.6 is - 5× cheaper output than Sol - 8.3× cheaper output than Fable - 2.5× cheaper input than Sol - 5× cheaper input than Fable to put this into perspective, $100 of grok output would cost $500 on sol
  • @kimmonismus @kimmonismus on x
    I'm taking this seriously now. Grok 4.6 was the leap I'd been hoping for. If the 10t model is still to come, then Elon's words can be taken seriously: it really could become the best model (in general). Although, of course, Anthropic already has Fable 5.5 ready and just waiting […
  • @mercor_ai @mercor_ai on x
    Grok 4.6 is one of the most cost efficient frontier models we've ever tested on APEX.
  • @synthwavedd Leo on x
    Holy shit lmao. Grok 4.6 benchmarks are out, and it's near SoTA on AA's Intelligence Index (61, matching Sol, 1 point behind Fable 5 and 2 points behind Opus 5), SoTA on GDPVal-AA v2, SoTA on AA-Briefcase and Harvey LAB (Vals), and 2nd only to Fable 5 on CursorBench v3.2, [image]
  • @aaronburnett Aaron Burnett on x
    I think it's possible that between Grok Bot and Grok 4.6 we start to see early signs of inflection from the SpaceXAI team. Which just a couple months ago was casually left off most people's leading AI Labs list. They may regret letting the rocket scientists have GPUs.
  • @luke_metro @luke_metro on x
    @tekbog This is how I feel about every fast follow model release, like codex/claude are so sticky Grok is funnier though due to all the Nazi and porn scandals
  • @mntruell Michael Truell on x
    Excited to release Grok 4.6. With each release, Grok is becoming a more capable digital colleague. 4.6 is significantly better at difficult tasks and knowledge work. It combines Opus-class intelligence and polish with very low cost and high speed.
  • @mattshumer_ Matt Shumer on x
    Looks like @SpaceXAI is catching up to the frontier. If this model is as good as the benchmarks say, they're on track to be a leading lab very soon... Super exciting shake-up in the AI race!
  • @benhylak Ben Hylak on x
    oh wow grok 4.6 beats fable! *looks closer*
  • @elonmusk Elon Musk on x
    Try out Grok 4.6! Double your tokens for next 7 days.
  • @dee_bosa Deirdre Bosa on x
    What a difference a year makes A year ago, frontier basically meant the big three US closed labs - OpenAI, Anthropic and Google. Now a credible list includes xAI and multiple Chinese/open weight labs.
  • @davis7 Ben Davis on x
    The Cursor + SpaceXAI comeback is glorious to watch - Grok Build is excellent, probably the best TUI from a lab (at least tied with Claude Code, Pi still clears all of them lol) - Grok Bot is the first novel take on non-coding agents since hermes/open claw and so far I'm really
  • @chrisgpt Chris on x
    Grok 4.6 is out!! and this is a much bigger jump over 4.5 It ties GPT 5.6 Sol at 61 on the Artificial Analysis Intelligence Index, leads GDPval-AA v2 and AA-Briefcase, and jumps from 15.7% to 26% on the new Terminal Bench V3. Fable 5 and 5.6 are still #1 on most coding [image]
  • @brendanfoody Brendan on x
    Grok 4.6 has massive improvements on both APEX-Agents and APEX-SWE. It's at the Pareto frontier of price and performance. Massive congratulations to @elonmusk and the @xai team.
  • @eliebakouch Elie on x
    interesting results from grok 4.6 model card: 1.  huge progress on DeepSearchQA and their internal KernelBench 2. grok models are the best on internal spaceXAI engineer benchmark (they likely trained on similar data but still impressive) 3. grok 4.6 also sota at inferenceEval whi…
  • @artificialanlys @artificialanlys on x
    Grok 4.6 is highly cost-effective: with headline pricing unchanged from Grok 4.5, it delivers a 5 point Intelligence Index gain at a cost per task comparable to Kimi K3 and far below Claude Opus 5, GPT-5.6 Sol, and Claude Fable 5. This places Grok 4.6 firmly on the Intelligence […
  • @cremieuxrecueil @cremieuxrecueil on x
    Grok? You mean the frontier model that's cheaper than K3 and more powerful than Fable?
  • @bindureddy Bindu Reddy on x
    Grok 4.6 just dropped and will be on LiveBench shortly - same price as Grok 4.5 which is great Not sure how much it really improves on 4.5. We will see
  • @milichab Andrew Milich on x
    Grok 4.6 is a leap in intelligence and persistence. From completing debugging tasks to building apps from scratch, Grok 4.6 is a fast, dependable, and smart companion. Try it now in Grok Build, Cursor, Grok Bot, and the SpaceXAI API!
  • @cremieuxrecueil @cremieuxrecueil on x
    Just finished my evaluation of it: - Bad at bioinformatics, behind Sol substantially - Better at making diagrams and making images publication-ready than K3/Sol - The best at writing proofs that are compact (Grok 4.5 was the previous best) - Still can't do CUDA programming well
  • @nvidia @nvidia on x
    Congrats to the @SpaceXAI team on the release of Grok 4.6. Grok 4.6 brings frontier intelligence, running and trained on NVIDIA GB300 NVL72 with NVLink to deliver exceptional performance, reliability and lowest token cost.
  • @emostaque Emad on x
    imo GDPVal is probably the most important benchmark, it measures the performance of models on real world tasks Big leap in performance here to top it at a great price, congrats to @SpaceXAI team & looking forward to the even bigger releases to come! https://openai.com/... [image]
  • @eliebakouch Elie on x
    grok 4.6 has a very impressive test time curve on cursorbench, in a league of its own they did more mid/pre-training on the grok 4.5 checkpoint and newer SFT stages with grok 4.5 traces and model base filtering. once again shows how important a good SFT ckpt is [image]
  • @elonmusk Elon Musk on x
    Grok 4.6 is a banger
  • @benjitaylor Benji Taylor on x
    The Grokening
  • @deryatr_ Derya Unutmaz on x
    I created this simulation app for engineering immune cells to fight cancer using Grok 4.6 build mode on my iPhone Grok app in just minutes! It simulates different engineering designs for cancer response and was created one-shot from a single-sentence prompt. This is pretty crazy!…
  • @chetanp Chetan Puttagunta on x
    Incredible achievement by SpaceXAI and Cursor. Enterprises and startups have already figured out Grok is a great model. The market opportunity is immense as Anthropic continues to be difficult to work with and is deeply unreliable as a partner/ vendor.
  • @kimmonismus @kimmonismus on x
    You know whats crazy? Grok 4.6 is not only cheaper than Opus 5 and 5.6 Sol. Its even cheaper than Sonnet 5. While at the same time, at least on par with the sota models. *Thats* the real moat. [image]
  • @deryatr_ Derya Unutmaz on x
    Testing Grok 4.6 now in coding and biomedical research; very preliminary outputs looking amazing! I will share specific examples soon. As I had predicted, Grok 4.6 has now become one of the few at the forefront of frontier AI models!
  • @mweinbach Max Weinbach on x
    Grok 4.6 seems pretty great on DeepSWE and GDPVal. Same price as Grok 4.5 too [image]
  • @ai_for_success AshutoshShrivastava on x
    My first test with Grok 4.6 I asked it to understand my current game code for Aether Ascent, the game I've been building for my son and many of you have seen, and then build a completely new stage. It created a new stage “10” called Stormwake. Love the stage design, especially [v…
  • @elonmusk Elon Musk on x
    @beffjezos Grok 4.7 is significantly better than 4.6 and should be ready in 3 to 4 weeks. Initial training is complete and now we're adding a massive amount of SpaceX company data in supplemental training. This will be something special.
  • r/SpaceXBets r on reddit
    If you wanna know why the stock just jumped:
  • r/grok r on reddit
    Grok 4.6 released