/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

← → days · ↑ ↓ browse · Enter similar · o open

Muse Spark 1.3 with max reasoning, in limited preview for partners, scores 62 on the Artificial Analysis Intelligence Index, behind only Fable 5.1 and Opus 5

@artificialanlys:

@artificialanlys

Context & Ripple Effects

Meta’s Muse Spark progression has moved quickly: the Muse Spark 1.1 public API preview emphasized advanced coding in July, and Muse Spark 1.2 reached 54 on the Artificial Analysis index in August. Version 1.3 raises that measured score to 62, placing it behind only Fable 5.1 and Opus 5 on the index.

The advance is not yet a broad release. Meta is limiting the max-reasoning configuration to partners, while its earlier coverage positioned Muse Spark as a model already powering Meta AI queries and shopping-mode queries.

First-order effects

  • Partners in the limited preview gain access to Meta’s highest-scoring disclosed Muse Spark configuration, while other developers remain outside that max-reasoning release.
  • Meta moves from the August tie for third at 54 to a 62 score behind only Fable 5.1 and Opus 5 on the Artificial Analysis index.

Second-order effects

  • Fable 5.1 and Opus 5 face a closer benchmark challenger in partner evaluations, making reasoning quality a more immediate comparison point for buyers with preview access.
  • Meta can use partner feedback from the restricted rollout to test the max-reasoning tier before extending availability beyond the earlier public API preview.

Third-order effects

  • The split between a publicly previewed Muse Spark 1.1 and partner-only max reasoning points to model vendors segmenting access by capability, with frontier reasoning becoming a controlled commercial tier rather than a uniform API feature.
  • If benchmark gains continue to arrive through limited previews, independent indexes will increasingly shape which models enterprise buyers seek to evaluate, even before general availability.

The trend: Frontier AI competition is shifting toward tiered reasoning models, where benchmark performance and controlled partner access jointly determine early market position.

Discussion

  • @kimmonismus @kimmonismus on x
    Today meta chose war with google. But hey, let them fight it out. That just means development will accelerate even further. By the way: I can't imagine OpenAI waiting long to reclaim the top spot in DeepSWE.
  • @mattdeitke Matt Deitke on x
    The jump for Muse Spark 1.3 on AA Index is quite strong! Only behind recent Anthropic models at this point. Much stronger models coming soon! 🍉🥳
  • @xeophon Florian Brand on x
    Gemini 3.8 held a spot at the pareto frontier for *checks notes* 3.5 hours
  • @zephyr_z9 @zephyr_z9 on x
    Very impressive
  • @angaisb_ Angel on x
    Not Meta mogging Gemini 3.8 Flash the same day lmao [image]
  • @altryne Alex Volkov on x
    Damn, gloves are off! Unreleased Meta Muse Spark with max reasoning is beating even Fable 5, while the xhigh that is available, matches Grok 4.6 and GPT 5.6 sol on @ArtificialAnlys This is quite the statement from @AIatMeta 🔥 Busy weeks ahead of us!
  • @bnjmn_marie Benjamin Marie on x
    Google was ahead only a few hours. And Meta will release the weights.
  • @haider1 Haider on x
    how tf Meta is moving this insanely fast needs to be studied they went from looking weirdly behind in the AI race to suddenly shipping at a pace that feels completely different
  • @chrisgpt Chris on x
    We're starting to see the fruits of Metas massive spend. So happy to see FAIR get its footing.
  • @spac89 @spac89 on x
    How the hell is Muse Spark 1.3 Max better than Fable 5? I genuinely don't think anyone saw this coming..
  • @udiwertheimer Udi Wertheimer on x
    what is even the point of anthropic anymore
  • @artificialanlys @artificialanlys on x
    Muse Spark 1.3 (max) scores 52% on Tau3-Bench Banking, the new #1 performer on this evaluation. Muse Spark 1.3 (xhigh) scores 47%, level with Claude Fable 5.1 (max, 47%) and GLM-5.3-Flash (47%) and behind its sibling in addition to e.g. Qwen3.8 Max (51%), Grok 4.6 (high, 51%), an…
  • @artificialanlys @artificialanlys on x
    The gains for Muse Spark 1.3 (xhigh) and Muse Spark 1.3 (max) over Muse Spark 1.2 on the Artificial Analysis Intelligence Index are concentrated in agentic evaluations: GDPval-AA v2 +94 and +139 Elo (1615 to 1709 and 1754), Terminal-Bench 2.1 +5 and +6 points (80% to 85% and 86%)…
  • @artificialanlys @artificialanlys on x
    Muse Spark 1.3 (xhigh) is the most cost-efficient model at its intelligence level: $0.55 per Intelligence Index task at Meta's unchanged $1.25/$4.25 per 1M token pricing. No model scoring 59 or above costs less per task. The nearest are Gemini 3.8 Flash (high, 59, $0.58), GPT-5.6…
  • @hesamation @hesamation on x
    Muse Spark 1.3 is now the top non-Anthropic model on Artificial Analysis Intelligence Index: > same score as Fable 5 > with 12x cheaper output tokens ($4.25/M vs $50/M) > just 1 point behind Opus 5, 4 points behind Fable 5.1 from YESTERDAY.  I'm especially curious about cost per …
  • @nateberkopec Nate Berkopec on x
    Muse Spark 1.3 xhigh is frontier for time/task AND cost/task. We have a new frontier human-in-the-loop coding model. Gemini 3.8-flash is slightly cheaper but also slower than 3.7, so no changes there.
  • @kimmonismus @kimmonismus on x
    It's worth really grasping this: in a very short time, the race between OpenAI and Anthropic has turned into a contest involving OpenAI, Anthropic, xAI, and Meta, and Google is back in the mix, too.  They are all on (mostly) equal footing, with little difference between them even…