/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

Artificial Analysis Coding Agent Index: GPT-6 Astra scored 67, roughly equal to Claude Opus 5, Fable 5, and Muse Spark 1.3, but trailing leader Fable 5.1's 70

Fable 5.1 in Claude Code leads the Index with a score of 70. … Various effort levels of the model occupy the Pareto frontier of token efficiency.

Artificial Analysis

Context & Ripple Effects

The coding-agent ranking arrives as the frontier has broadened beyond a single vendor: Muse Spark 1.3 entered limited partner preview with a 62 Intelligence Index score, while Fable 5.1 and Opus 5 led that separate measure. GPT-6 Astra places into a tightly packed coding-agent tier rather than establishing a clear lead.

Artificial Analysis also frames Astra's different effort levels as a token-efficiency Pareto frontier. That makes the comparison relevant to buyers selecting an agent configuration, not just a model name.

First-order effects

  • Fable 5.1 retains the Coding Agent Index lead at 70, while GPT-6 Astra joins Claude Opus 5, Fable 5, and Muse Spark 1.3 at 67, giving coding-agent buyers several near-frontier options.
  • Astra's effort-level configurations give users a direct trade-off between token use and coding-agent performance, according to the index's frontier analysis.

Second-order effects

  • Fable, Claude, and Muse face pressure to demonstrate task-level efficiency alongside raw coding scores, because a near-leading model with lower token consumption can alter deployment costs.
  • Teams evaluating coding agents will need to compare model-and-effort configurations rather than rely on a single per-token price or headline benchmark rank.

Third-order effects

  • If frontier models continue to cluster on agent benchmarks, differentiation shifts toward availability, deployment access, and inference efficiency rather than a durable capability lead.
  • Coding-agent procurement is likely to center on cost per completed task, with benchmark leaders needing to show that their performance premium justifies their inference budget.

The trend: Frontier coding agents are converging on benchmark capability while competition moves toward token-efficient reasoning configurations and task-level economics.

Discussion

  • @artificialanlys @artificialanlys on x
    GPT-6 Astra makes significant gains in the Artificial Analysis Coding Agent Index, scoring equal to Fable 5 at lower cost. In the Intelligence Index, it uses fewer tokens than GPT-5.6 Sol for similar performance, but this is outweighed by higher prices Pricing is 2.5x GPT-5.6 Sol…
  • @theo @theo on x
    Oh my god, Gemini 3.8 Flash beat Astra on DeepSWE 💀 73.8% vs 73.3%
  • @cheatyyyy @cheatyyyy on x
    GPT-6 Astra is on the pareto-frontier of cost efficiency due to being EXTREMELY token efficient. It is in a whole league of it's own. It is cheaper than Gemini 3.8 Flash per task (a model 13x cheaper than Astra).
  • @robdonnelly47 Rob Donnelly on x
    @cheatyyyy The way the cost curve bends backwards on Arc AGI blows my mind. The Max reasoning level is so much better at solving the problems quickly that it actually costs less than if you tried with a lower reasoning effort level
  • @stevenheidel Steven Heidel on x
    token pricing is effectively meaningless now. 3.8 flash looks 13x cheaper when measured per token, but Astra is cheaper per task since it's far more efficient. measure your costs per task, not per token.
  • @seanjtaylor Sean J. Taylor on x
    Everyone's gotta have an Astra take, so I'll give you mine: we are making fast progress on eradicating hallucinations. This eval is not some academic benchmark; it very realistically captures user experience. https://deploymentsafety.openai.com/ ...
  • @stevehou Steve Hou on x
    Dollar per task is dropping precipitously. I think you can be bearish on AI as an investment (cycle), and still marvel and celebrate the extraordinary technological progress made by/in artificial intelligence. The trouble with some of the AI bears is that they are so committed to…
  • @alexandr_wang Alexandr Wang on x
    1/ today we're releasing muse spark 1.3—available in muse code & the meta model api. this is our most capable model yet—frontier performance almost too cheap to meter. much stronger at agentic and coding with better usability. we think users will really notice the jump.
  • @ashvinair Ashvin Nair on x
    I recently joined the science of reasoning team at Meta! AI progress is moving faster than anyone can adjust to. I'm finding the goal of bringing personal superintelligence to everyone — AI that's positive and useful in people's lives worldwide — to be very motivating. Go try Mus…
  • @dryangsong Yang Song on x
    Muse Spark 1.3 is here! The leap from 1.0 to 1.3 is nothing short of a miracle, and it took a lot of magic to make it happen. So proud of our team of magicians who sparked this incredible leap. And we are just getting started!
  • @alexandr_wang Alexandr Wang on x
    for a single dime look at what muse spark can give you
  • @quinnypig Corey Quinn on x
    Meta casting shade at Google in AI is a playground slapfight outside a MMA championship.
  • @alexandr_wang Alexandr Wang on x
    muse spark 1.3 + muse code evals competitively with Claude code + opus 5 and Claude code + fable 5
  • @altryne Alex Volkov on x
    Damn, gloves are off! Unreleased Meta Muse Spark with max reasoning is beating even Fable 5, while the xhigh that is available, matches Grok 4.6 and GPT 5.6 sol on @ArtificialAnlys This is quite the statement from @AIatMeta 🔥 Busy weeks ahead of us!
  • @deryatr_ Derya Unutmaz on x
    I wanted to try something quick with Muse Spark 1.3, so I had it build this pirate platformer. It was incredibly fast and worked flawlessly at one-shot! Now I'm enhancing the artwork, sprites and adding new levels 😊This is such a fun coding model, it just works, fast and cheap!
  • @xeophon Florian Brand on x
    Gemini 3.8 held a spot at the pareto frontier for *checks notes* 3.5 hours
  • @bnjmn_marie Benjamin Marie on x
    Google was ahead only a few hours. And Meta will release the weights.
  • @alexandr_wang Alexandr Wang on x
    i really hate to say it, but... gemini who? 🏎️💨
  • @hesamation @hesamation on x
    Muse Spark 1.3 is now the top non-Anthropic model on Artificial Analysis Intelligence Index: > same score as Fable 5 > with 12x cheaper output tokens ($4.25/M vs $50/M) > just 1 point behind Opus 5, 4 points behind Fable 5.1 from YESTERDAY.  I'm especially curious about cost per …
  • @metafordevs @metafordevs on x
    Muse Spark 1.3 is now available in Muse Code and Meta Model API. It's tuned for the agentic builds developers actually ship, including long-running, multi-agent workflows. 🧵👇(1/4) [video]
  • @alexandr_wang Alexandr Wang on x
    underrated result of muse spark 1.3 try it out with muse code: curl -fsSL https://dev.meta.ai/... | bash
  • @aiatmeta @aiatmeta on x
    We're excited to release Muse Spark 1.3 with improved performance on agentic and coding tasks, and a focus on real-world usability. Key capabilities: → Sustains longer-horizon work across multiple workflows in a single thread → More actively collaborates with users: it asks clari…
  • @edwardsun0909 Zhiqing Sun on x
    We posttrained avocado again and enabled better test-time scaling. It's significantly better than its predecessor in all capabilities!
  • @dandr1s Dan on x
    Meta's Muse Spark 1.3 just scored 62 on Artificial Analysis' Intelligence Index. That ties Claude Fable 5, but Muse costs 8x less for input and nearly 12x less for output. It also scores above GPT-5.6 Sol, Grok 4.6, Kimi K3, and Gemini 3.8 Flash. Meta is suddenly in the top tier …
  • @openrouter @openrouter on x
    Muse Spark 1.3 is now available on OpenRouter! Built for long-running agentic, multi-agent, and coding workflows. It tracks what it learns, handles conflicting inputs, and asks for clarification or confirmation when needed. Use it now: https://openrouter.ai/...
  • @mihawkxxxxx @mihawkxxxxx on x
    @DanDr1s Incredible Minecraft from muse spark 1.3 cost 10 cents
  • @nateberkopec Nate Berkopec on x
    Muse Spark 1.3 xhigh is frontier for time/task AND cost/task. We have a new frontier human-in-the-loop coding model. Gemini 3.8-flash is slightly cheaper but also slower than 3.7, so no changes there.
  • @kimmonismus @kimmonismus on x
    It's worth really grasping this: in a very short time, the race between OpenAI and Anthropic has turned into a contest involving OpenAI, Anthropic, xAI, and Meta, and Google is back in the mix, too.  They are all on (mostly) equal footing, with little difference between them even…
  • @kimmonismus @kimmonismus on x
    Today meta chose war with google. But hey, let them fight it out. That just means development will accelerate even further. By the way: I can't imagine OpenAI waiting long to reclaim the top spot in DeepSWE.
  • @mattdeitke Matt Deitke on x
    The jump for Muse Spark 1.3 on AA Index is quite strong! Only behind recent Anthropic models at this point. Much stronger models coming soon! 🍉🥳
  • @zephyr_z9 @zephyr_z9 on x
    Very impressive
  • @angaisb_ Angel on x
    Not Meta mogging Gemini 3.8 Flash the same day lmao [image]
  • @haider1 Haider on x
    how tf Meta is moving this insanely fast needs to be studied they went from looking weirdly behind in the AI race to suddenly shipping at a pace that feels completely different
  • @chrisgpt Chris on x
    We're starting to see the fruits of Metas massive spend. So happy to see FAIR get its footing.
  • @spac89 @spac89 on x
    How the hell is Muse Spark 1.3 Max better than Fable 5? I genuinely don't think anyone saw this coming..
  • @udiwertheimer Udi Wertheimer on x
    what is even the point of anthropic anymore
  • @artificialanlys @artificialanlys on x
    Muse Spark 1.3 (max) scores 52% on Tau3-Bench Banking, the new #1 performer on this evaluation. Muse Spark 1.3 (xhigh) scores 47%, level with Claude Fable 5.1 (max, 47%) and GLM-5.3-Flash (47%) and behind its sibling in addition to e.g. Qwen3.8 Max (51%), Grok 4.6 (high, 51%), an…
  • @artificialanlys @artificialanlys on x
    The gains for Muse Spark 1.3 (xhigh) and Muse Spark 1.3 (max) over Muse Spark 1.2 on the Artificial Analysis Intelligence Index are concentrated in agentic evaluations: GDPval-AA v2 +94 and +139 Elo (1615 to 1709 and 1754), Terminal-Bench 2.1 +5 and +6 points (80% to 85% and 86%)…
  • @artificialanlys @artificialanlys on x
    Muse Spark 1.3 (xhigh) is the most cost-efficient model at its intelligence level: $0.55 per Intelligence Index task at Meta's unchanged $1.25/$4.25 per 1M token pricing.  No model scoring 59 or above costs less per task.  The nearest are Gemini 3.8 Flash (high, 59, $0.58), GPT-5…
  • @rihardjarc Rihard Jarc on x
    Wow, $META full Scorched-Earth Strategy with this AI model pricing. Nobody wants to have Zuck as a competitor.
  • @finkd Mark Zuckerberg on x
    Mark Zuckerberg says Meta's Watermelon model and Muse Spark open weights are “coming soon”
  • Brian Gamido Brian Gamido on linkedin
    We just released Muse Spark 1.3.  —  It's Meta's biggest jump so far on AI performance, especially on agentic and coding tasks …
  • @hesamation @hesamation on x
    some notes on Astra's surprisingly low (61) intelligence score on Artificial Analysis: it seems like Astra is extremely spiky in benchmarks, crushes some, and regresses on others. > OpenAI's benchmarks: Astra crushes Fable 5.1 and spits on it. > AA intelligence score: incredibly …
  • @bindureddy Bindu Reddy on x
    🚨 GPT 6-Astra Is AWESOME, But Is Slightly Below Fable 5.1 - much cheaper and faster than Fable 5.1 - worse on long-running agentic loops - much better at browser use and 3D renders OpenAI is back in the game!