Artificial Analysis Coding Agent Index: GPT-6 Astra scored 67, roughly equal to Claude Opus 5, Fable 5, and Muse Spark 1.3, but trailing leader Fable 5.1's 70
Fable 5.1 in Claude Code leads the Index with a score of 70. … Various effort levels of the model occupy the Pareto frontier of token efficiency.
Artificial Analysis
Context & Ripple Effects
The coding-agent ranking arrives as the frontier has broadened beyond a single vendor: Muse Spark 1.3 entered limited partner preview with a 62 Intelligence Index score, while Fable 5.1 and Opus 5 led that separate measure. GPT-6 Astra places into a tightly packed coding-agent tier rather than establishing a clear lead.
Artificial Analysis also frames Astra's different effort levels as a token-efficiency Pareto frontier. That makes the comparison relevant to buyers selecting an agent configuration, not just a model name.
First-order effects
- Fable 5.1 retains the Coding Agent Index lead at 70, while GPT-6 Astra joins Claude Opus 5, Fable 5, and Muse Spark 1.3 at 67, giving coding-agent buyers several near-frontier options.
- Astra's effort-level configurations give users a direct trade-off between token use and coding-agent performance, according to the index's frontier analysis.
Second-order effects
- Fable, Claude, and Muse face pressure to demonstrate task-level efficiency alongside raw coding scores, because a near-leading model with lower token consumption can alter deployment costs.
- Teams evaluating coding agents will need to compare model-and-effort configurations rather than rely on a single per-token price or headline benchmark rank.
Third-order effects
- If frontier models continue to cluster on agent benchmarks, differentiation shifts toward availability, deployment access, and inference efficiency rather than a durable capability lead.
- Coding-agent procurement is likely to center on cost per completed task, with benchmark leaders needing to show that their performance premium justifies their inference budget.
The trend: Frontier coding agents are converging on benchmark capability while competition moves toward token-efficient reasoning configurations and task-level economics.
Related: Agent Inference Economics · AI cost per useful task · Claude Opus 5 · Fable · Muse Spark 1.3 reaches the intelligence frontier
Related Coverage
- GPT-6 Astra — It's going to be API priced at the same rate as Claude Fable 5 and 5.1: $10/million input and $50/million output. Simon Willison's Weblog · Simon Willison
- OpenAI's “generational leap” with GPT-6 Astra The Rundown AI
- GPT-6 Can Downplay Its Own Abilities in Tests Through „Sandbagging" Trending Topics · Jakob Steinschaden
- GPT-6 Astra makes major gains in the Artificial Analysis Coding Agent Index Hacker News
- Elevated errors for multiple models Claude
- Muse Spark 1.3: Meta reaches the frontier Artificial Analysis
- Introducing Muse Spark 1.3 Meta AI Research
- GPT-6 Astra Trails Top Models From Anthropic and Meta in Benchmarks Trending Topics · Jakob Steinschaden
- Benchmarks disagree on GPT-6 Astra, but its human-beating efficiency on ARC-AGI-3 pulls Chollet's AGI forecast forward The Decoder · Maximilian Schreiner
- OpenAI will sell you Astra, but not the system that scored 98.6% on ARC-AGI-3 The New Stack · Matthew Burns
- GPT-6 Astra Arrives With Major Gains, Staged Access, and New Questions About Its Benchmarks WinBuzzer · Markus Kasanmascheff
Discussion
-
@artificialanlys
@artificialanlys
on x
GPT-6 Astra makes significant gains in the Artificial Analysis Coding Agent Index, scoring equal to Fable 5 at lower cost. In the Intelligence Index, it uses fewer tokens than GPT-5.6 Sol for similar performance, but this is outweighed by higher prices Pricing is 2.5x GPT-5.6 Sol…
-
@theo
@theo
on x
Oh my god, Gemini 3.8 Flash beat Astra on DeepSWE 💀 73.8% vs 73.3%
-
@cheatyyyy
@cheatyyyy
on x
GPT-6 Astra is on the pareto-frontier of cost efficiency due to being EXTREMELY token efficient. It is in a whole league of it's own. It is cheaper than Gemini 3.8 Flash per task (a model 13x cheaper than Astra).
-
@robdonnelly47
Rob Donnelly
on x
@cheatyyyy The way the cost curve bends backwards on Arc AGI blows my mind. The Max reasoning level is so much better at solving the problems quickly that it actually costs less than if you tried with a lower reasoning effort level
-
@stevenheidel
Steven Heidel
on x
token pricing is effectively meaningless now. 3.8 flash looks 13x cheaper when measured per token, but Astra is cheaper per task since it's far more efficient. measure your costs per task, not per token.
-
@seanjtaylor
Sean J. Taylor
on x
Everyone's gotta have an Astra take, so I'll give you mine: we are making fast progress on eradicating hallucinations. This eval is not some academic benchmark; it very realistically captures user experience. https://deploymentsafety.openai.com/ ...
-
@stevehou
Steve Hou
on x
Dollar per task is dropping precipitously. I think you can be bearish on AI as an investment (cycle), and still marvel and celebrate the extraordinary technological progress made by/in artificial intelligence. The trouble with some of the AI bears is that they are so committed to…
-
@alexandr_wang
Alexandr Wang
on x
1/ today we're releasing muse spark 1.3—available in muse code & the meta model api. this is our most capable model yet—frontier performance almost too cheap to meter. much stronger at agentic and coding with better usability. we think users will really notice the jump.
-
@ashvinair
Ashvin Nair
on x
I recently joined the science of reasoning team at Meta! AI progress is moving faster than anyone can adjust to. I'm finding the goal of bringing personal superintelligence to everyone — AI that's positive and useful in people's lives worldwide — to be very motivating. Go try Mus…
-
@dryangsong
Yang Song
on x
Muse Spark 1.3 is here! The leap from 1.0 to 1.3 is nothing short of a miracle, and it took a lot of magic to make it happen. So proud of our team of magicians who sparked this incredible leap. And we are just getting started!
-
@alexandr_wang
Alexandr Wang
on x
for a single dime look at what muse spark can give you
-
@quinnypig
Corey Quinn
on x
Meta casting shade at Google in AI is a playground slapfight outside a MMA championship.
-
@alexandr_wang
Alexandr Wang
on x
muse spark 1.3 + muse code evals competitively with Claude code + opus 5 and Claude code + fable 5
-
@altryne
Alex Volkov
on x
Damn, gloves are off! Unreleased Meta Muse Spark with max reasoning is beating even Fable 5, while the xhigh that is available, matches Grok 4.6 and GPT 5.6 sol on @ArtificialAnlys This is quite the statement from @AIatMeta 🔥 Busy weeks ahead of us!
-
@deryatr_
Derya Unutmaz
on x
I wanted to try something quick with Muse Spark 1.3, so I had it build this pirate platformer. It was incredibly fast and worked flawlessly at one-shot! Now I'm enhancing the artwork, sprites and adding new levels 😊This is such a fun coding model, it just works, fast and cheap!
-
@xeophon
Florian Brand
on x
Gemini 3.8 held a spot at the pareto frontier for *checks notes* 3.5 hours
-
@bnjmn_marie
Benjamin Marie
on x
Google was ahead only a few hours. And Meta will release the weights.
-
@alexandr_wang
Alexandr Wang
on x
i really hate to say it, but... gemini who? 🏎️💨
-
@hesamation
@hesamation
on x
Muse Spark 1.3 is now the top non-Anthropic model on Artificial Analysis Intelligence Index: > same score as Fable 5 > with 12x cheaper output tokens ($4.25/M vs $50/M) > just 1 point behind Opus 5, 4 points behind Fable 5.1 from YESTERDAY. I'm especially curious about cost per …
-
@metafordevs
@metafordevs
on x
Muse Spark 1.3 is now available in Muse Code and Meta Model API. It's tuned for the agentic builds developers actually ship, including long-running, multi-agent workflows. 🧵👇(1/4) [video]
-
@alexandr_wang
Alexandr Wang
on x
underrated result of muse spark 1.3 try it out with muse code: curl -fsSL https://dev.meta.ai/... | bash
-
@aiatmeta
@aiatmeta
on x
We're excited to release Muse Spark 1.3 with improved performance on agentic and coding tasks, and a focus on real-world usability. Key capabilities: → Sustains longer-horizon work across multiple workflows in a single thread → More actively collaborates with users: it asks clari…
-
@edwardsun0909
Zhiqing Sun
on x
We posttrained avocado again and enabled better test-time scaling. It's significantly better than its predecessor in all capabilities!
-
@dandr1s
Dan
on x
Meta's Muse Spark 1.3 just scored 62 on Artificial Analysis' Intelligence Index. That ties Claude Fable 5, but Muse costs 8x less for input and nearly 12x less for output. It also scores above GPT-5.6 Sol, Grok 4.6, Kimi K3, and Gemini 3.8 Flash. Meta is suddenly in the top tier …
-
@openrouter
@openrouter
on x
Muse Spark 1.3 is now available on OpenRouter! Built for long-running agentic, multi-agent, and coding workflows. It tracks what it learns, handles conflicting inputs, and asks for clarification or confirmation when needed. Use it now: https://openrouter.ai/...
-
@mihawkxxxxx
@mihawkxxxxx
on x
@DanDr1s Incredible Minecraft from muse spark 1.3 cost 10 cents
-
@nateberkopec
Nate Berkopec
on x
Muse Spark 1.3 xhigh is frontier for time/task AND cost/task. We have a new frontier human-in-the-loop coding model. Gemini 3.8-flash is slightly cheaper but also slower than 3.7, so no changes there.
-
@kimmonismus
@kimmonismus
on x
It's worth really grasping this: in a very short time, the race between OpenAI and Anthropic has turned into a contest involving OpenAI, Anthropic, xAI, and Meta, and Google is back in the mix, too. They are all on (mostly) equal footing, with little difference between them even…
-
@kimmonismus
@kimmonismus
on x
Today meta chose war with google. But hey, let them fight it out. That just means development will accelerate even further. By the way: I can't imagine OpenAI waiting long to reclaim the top spot in DeepSWE.
-
@mattdeitke
Matt Deitke
on x
The jump for Muse Spark 1.3 on AA Index is quite strong! Only behind recent Anthropic models at this point. Much stronger models coming soon! 🍉🥳
-
@zephyr_z9
@zephyr_z9
on x
Very impressive
-
@angaisb_
Angel
on x
Not Meta mogging Gemini 3.8 Flash the same day lmao [image]
-
@haider1
Haider
on x
how tf Meta is moving this insanely fast needs to be studied they went from looking weirdly behind in the AI race to suddenly shipping at a pace that feels completely different
-
@chrisgpt
Chris
on x
We're starting to see the fruits of Metas massive spend. So happy to see FAIR get its footing.
-
@spac89
@spac89
on x
How the hell is Muse Spark 1.3 Max better than Fable 5? I genuinely don't think anyone saw this coming..
-
@udiwertheimer
Udi Wertheimer
on x
what is even the point of anthropic anymore
-
@artificialanlys
@artificialanlys
on x
Muse Spark 1.3 (max) scores 52% on Tau3-Bench Banking, the new #1 performer on this evaluation. Muse Spark 1.3 (xhigh) scores 47%, level with Claude Fable 5.1 (max, 47%) and GLM-5.3-Flash (47%) and behind its sibling in addition to e.g. Qwen3.8 Max (51%), Grok 4.6 (high, 51%), an…
-
@artificialanlys
@artificialanlys
on x
The gains for Muse Spark 1.3 (xhigh) and Muse Spark 1.3 (max) over Muse Spark 1.2 on the Artificial Analysis Intelligence Index are concentrated in agentic evaluations: GDPval-AA v2 +94 and +139 Elo (1615 to 1709 and 1754), Terminal-Bench 2.1 +5 and +6 points (80% to 85% and 86%)…
-
@artificialanlys
@artificialanlys
on x
Muse Spark 1.3 (xhigh) is the most cost-efficient model at its intelligence level: $0.55 per Intelligence Index task at Meta's unchanged $1.25/$4.25 per 1M token pricing. No model scoring 59 or above costs less per task. The nearest are Gemini 3.8 Flash (high, 59, $0.58), GPT-5…
-
@rihardjarc
Rihard Jarc
on x
Wow, $META full Scorched-Earth Strategy with this AI model pricing. Nobody wants to have Zuck as a competitor.
-
@finkd
Mark Zuckerberg
on x
Mark Zuckerberg says Meta's Watermelon model and Muse Spark open weights are “coming soon”
-
Brian Gamido
Brian Gamido
on linkedin
We just released Muse Spark 1.3. — It's Meta's biggest jump so far on AI performance, especially on agentic and coding tasks …
-
@hesamation
@hesamation
on x
some notes on Astra's surprisingly low (61) intelligence score on Artificial Analysis: it seems like Astra is extremely spiky in benchmarks, crushes some, and regresses on others. > OpenAI's benchmarks: Astra crushes Fable 5.1 and spits on it. > AA intelligence score: incredibly …
-
@bindureddy
Bindu Reddy
on x
🚨 GPT 6-Astra Is AWESOME, But Is Slightly Below Fable 5.1 - much cheaper and faster than Fable 5.1 - worse on long-running agentic loops - much better at browser use and 3D renders OpenAI is back in the game!