Muse Spark 1.3 with max reasoning, in limited preview for partners, scores 62 on the Artificial Analysis Intelligence Index, behind only Fable 5.1 and Opus 5
@artificialanlys:
@artificialanlys
Context & Ripple Effects
Meta’s Muse Spark progression has moved quickly: the Muse Spark 1.1 public API preview emphasized advanced coding in July, and Muse Spark 1.2 reached 54 on the Artificial Analysis index in August. Version 1.3 raises that measured score to 62, placing it behind only Fable 5.1 and Opus 5 on the index.
The advance is not yet a broad release. Meta is limiting the max-reasoning configuration to partners, while its earlier coverage positioned Muse Spark as a model already powering Meta AI queries and shopping-mode queries.
First-order effects
Partners in the limited preview gain access to Meta’s highest-scoring disclosed Muse Spark configuration, while other developers remain outside that max-reasoning release.
Meta moves from the August tie for third at 54 to a 62 score behind only Fable 5.1 and Opus 5 on the Artificial Analysis index.
Second-order effects
Fable 5.1 and Opus 5 face a closer benchmark challenger in partner evaluations, making reasoning quality a more immediate comparison point for buyers with preview access.
Meta can use partner feedback from the restricted rollout to test the max-reasoning tier before extending availability beyond the earlier public API preview.
Third-order effects
The split between a publicly previewed Muse Spark 1.1 and partner-only max reasoning points to model vendors segmenting access by capability, with frontier reasoning becoming a controlled commercial tier rather than a uniform API feature.
If benchmark gains continue to arrive through limited previews, independent indexes will increasingly shape which models enterprise buyers seek to evaluate, even before general availability.
The trend: Frontier AI competition is shifting toward tiered reasoning models, where benchmark performance and controlled partner access jointly determine early market position.
Today meta chose war with google. But hey, let them fight it out. That just means development will accelerate even further. By the way: I can't imagine OpenAI waiting long to reclaim the top spot in DeepSWE.
Damn, gloves are off! Unreleased Meta Muse Spark with max reasoning is beating even Fable 5, while the xhigh that is available, matches Grok 4.6 and GPT 5.6 sol on @ArtificialAnlys This is quite the statement from @AIatMeta 🔥 Busy weeks ahead of us!
how tf Meta is moving this insanely fast needs to be studied they went from looking weirdly behind in the AI race to suddenly shipping at a pace that feels completely different
Muse Spark 1.3 (max) scores 52% on Tau3-Bench Banking, the new #1 performer on this evaluation. Muse Spark 1.3 (xhigh) scores 47%, level with Claude Fable 5.1 (max, 47%) and GLM-5.3-Flash (47%) and behind its sibling in addition to e.g. Qwen3.8 Max (51%), Grok 4.6 (high, 51%), an…
The gains for Muse Spark 1.3 (xhigh) and Muse Spark 1.3 (max) over Muse Spark 1.2 on the Artificial Analysis Intelligence Index are concentrated in agentic evaluations: GDPval-AA v2 +94 and +139 Elo (1615 to 1709 and 1754), Terminal-Bench 2.1 +5 and +6 points (80% to 85% and 86%)…
Muse Spark 1.3 (xhigh) is the most cost-efficient model at its intelligence level: $0.55 per Intelligence Index task at Meta's unchanged $1.25/$4.25 per 1M token pricing. No model scoring 59 or above costs less per task. The nearest are Gemini 3.8 Flash (high, 59, $0.58), GPT-5.6…
Muse Spark 1.3 is now the top non-Anthropic model on Artificial Analysis Intelligence Index: > same score as Fable 5 > with 12x cheaper output tokens ($4.25/M vs $50/M) > just 1 point behind Opus 5, 4 points behind Fable 5.1 from YESTERDAY. I'm especially curious about cost per …
Muse Spark 1.3 xhigh is frontier for time/task AND cost/task. We have a new frontier human-in-the-loop coding model. Gemini 3.8-flash is slightly cheaper but also slower than 3.7, so no changes there.
It's worth really grasping this: in a very short time, the race between OpenAI and Anthropic has turned into a contest involving OpenAI, Anthropic, xAI, and Meta, and Google is back in the mix, too. They are all on (mostly) equal footing, with little difference between them even…