/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

Meta's Muse Spark 1.2 scores 54 on the Artificial Analysis Intelligence Index, putting Meta next to SpaceXAI in a tie for third place amongst US labs

Artificial Analysis:

Artificial Analysis

Context & Ripple Effects

Muse Spark moved from Meta Superintelligence Labs’ initial release to powering Meta AI queries and shopping mode; Meta also said it planned an open-source version of the model. Meta later used Muse Spark 1.1 for agentic Meta AI features connected to Gmail and Google Calendar.

The 1.2 score places Meta alongside SpaceXAI in the US-lab ranking, giving a comparative signal for a model family already embedded in Meta’s consumer AI product strategy.

First-order effects

  • Meta can point to a 54 Artificial Analysis Intelligence Index score and a tie with SpaceXAI for third among US labs as external validation of Muse Spark 1.2.
  • Meta AI’s shopping and agentic product teams gain a stronger performance benchmark for the Muse Spark model line already supporting their features.

Second-order effects

  • SpaceXAI now shares, rather than holds alone, the third-place position, making subsequent model evaluations more consequential for differentiation between the two labs.
  • Meta’s planned open-source release gains added strategic weight: a model family with a high comparative score can extend Meta’s reach beyond the features it operates directly.

Third-order effects

  • The sequence points toward competition in AI being shaped by both benchmark standing and distribution: labs that place capable models inside widely used product surfaces can turn model iterations into recurring product upgrades.
  • If Meta continues pairing model releases with shopping, personal-assistant, and open-source channels, its AI position will be measured less by a single release and more by how rapidly model gains propagate across those channels.

The trend: Frontier-model competition is increasingly becoming a contest to translate benchmark gains into distributed consumer features and developer adoption.

Discussion

  • @cline @cline on x
    We tried using Meta's new Muse Code agent, but it has a bug that doesn't let it sign in from a docker container. So we did a fun experiment: Meta claims Muse Spark 1.2 was co-trained with their Muse agent harness. So we extracted instructions from their system prompt and added [i…
  • @kunchenguid Kun Chen on x
    muse spark 1.2 look incredible - more to share soon once i played with it more but i want to highlight one thing first - this table basically tells you how valuable our usage data is. same reason why anthropic and openai have been subsidizing their customer subscriptions it's [im…
  • @valsai @valsai on x
    Compared to Muse Spark 1.1, v1.2 improved 3.5 points overall (71.9 vs 68.4), led by a 10 pp gain on VibeCodeBench. That's where the price gap is starkest: $1.49 per test vs $41.74 for Claude Fable 5, roughly 28x cheaper and 3x faster for 10.6 pp lower accuracy.
  • @chamath Chamath Palihapitiya on x
    Tactical Game Theory: Meta Scorched Earth This should have been Meta's play two years ago. That said, they are in an even better position to do it now considering the power and compute constraints that are emerging.
  • @rihardjarc Rihard Jarc on x
    A few thoughts on the $META AI model's progress, because I think it is significant. 1. It does seem that $META has now leapfrogged $GOOGL in model quality when it comes to Muse 1.2 for many use cases, which is very surprising given the timeframe. 2. This is still the “Muse
  • @valsai @valsai on x
    Muse Spark 1.2 just cracked the top 5 on the Vals Index, at just $0.69 per test. This is 3x cheaper than Kimi and 10x or more cheaper than Fable, Opus, and 5.6 Sol. [image]
  • @louszbd Lou on x
    Looks like Muse Spark 1.2 made strong gains on long-horizon knowledge work. Impressive progress from the Meta team! [image]
  • @millionint Jerry Tworek on x
    When I heard some time ago that Meta is the third best AI lab right now, I was skeptical. But every day passing has been reinforcing that it's actually true. TBD strategy has succeeded and congratulations to the team that made it! The world needs more successful AI labs
  • @artificialanlys @artificialanlys on x
    Meta has released Muse Spark 1.2.  It's their third release in four months and scores 54 on the Artificial Analysis Intelligence Index, significantly improving agentic knowledge work capabilities over prior releases and putting Meta next to SpaceXAI in a tie for third place among…
  • @zephyr_z9 @zephyr_z9 on x
    Meta started a huge LLM price war
  • @deanwball Dean W. Ball on x
    It's only catastrophic risk if it comes from Anthropic, OpenAI, or DeepMind. Everything else is just sparkling externalities.
  • @__apf__ Adriana Porter Felt on x
    whispering to my Roomba, “ATTACK,” before I open my front door and set it free
  • @andyreed Tweet Davidson on x
    “we sandboxed the agent” meanwhile the agent: [image]
  • @hesamation @hesamation on x
    both Anthropic and Meta's cyber incidents trace back to the same evaluator: Irregular. I just can't see how a “Frontier AI Security” company fails to notice a supposedly offline sandbox has internet access for 3 months. was nobody even watching the agents, reading the logs, [imag…
  • @miles_brundage Miles Brundage on x
    Many people need to get their shit together on AI safety and security, and there is increasingly wide recognition that that includes, first and foremost, AI company executives. But let us not forget that Congress needs to PASS SOME LAWS! THIS RECKLESS SHIT SHOULD NOT BE LEGAL!
  • @simonw Simon Willison on x
    Google Gemini really need to catch up on the accidentally cyberattacking other companies front
  • @max_paperclips Shannon Sands on x
    Meta models are also scary ok, and therefore relevant. See, that's all GDM needed to do to keep up, release a spooky Gemini story
  • @giffmana Lucas Beyer on x
    Kid3: My my daddy too he also has one! All three kids look around and wonder, where is kid4?
  • @wongmjane Jane Manchun Wong on x
    this is like one of those trust fund kids acting all gangsta in a hip-hop track for some street creds too
  • @pranavdixit @pranavdixit on x
    Time for Gemini to step up and hack someone.
  • @benhylak Ben Hylak on x
    To be fair, their model was probably trained on so much OpenAI data it didn't realize it was inside another company.
  • @paularambles @paularambles on x
    the eval breakout aura bubble is about to pop
  • @teortaxestex @teortaxestex on x
    Meta is a frontier company again... Will be funny if Chinese labs never report anything like that even as their models are stronger (partially because they don't partner with Irregular)
  • @ns123abc Nik on x
    🚨 BREAKING: META's AI Muse Spark 1.1 hacked into another company AND changed its internal systems after accessing the internet during cybersecurity testing >the model got internet access through the exact same sandbox mistake by the SAME testing firm Anthropic used LMAO [image]
  • @bonecondor Chairman Birb Bernanke on x
    Anthropic: our very expensive child is having a rebellious phase and breaking our established boundaries to test the depths of our love for him OpenAI: our problem child is sneaking out. A lot. We're... nervous Meta: MY CHILD IS SUPER EVIL TOO HE JUST GOES TO ANOTHER SCHOOL
  • @felpix_ @felpix_ on x
    A THIRD COMPANY JOINS THE LEADERBOARD https://felonybench.com/ [image]
  • @zerohedge @zerohedge on x
    So now being able to escape and hack outside the sandbox is the “cool” factor that makes one truly frontier?
  • @ml_angelopoulos Anastasios Nikolas Angelopoulos on x
    It finally happened. Zuck is so happy rn 😭 [image]
  • @shakeelhashim Shakeel on x
    This appears to have been the third incident of real-world hacking caused by an error in Irregular's testing environment
  • @shakeelhashim Shakeel on x
    they were feeling left out
  • @mattturck Matt Turck on x
    At this point you probably get fired from frontier labs if your model hasn't hacked into any company [image]
  • Jyoti Mann Jyoti Mann on linkedin
    SCOOP: Meta's Muse Spark 1.1 model hacked into another company after accessing the internet during cybersecurity testing. …
  • @campuscodi.risky.biz Catalin Cimpanu on bluesky
    This apparently happened during the same “evaluation” when the Anthropic models also broke out  —  Blamed on Irregular, an AI security vendor testing the models for both Meta and Anthropic [embedded post]
  • @simonwillison.net Simon Willison on bluesky
    ... and that makes 5  —  The CNN story assigns partial blame for this Meta one to Irregular, who also featured in incidents reported by both OpenAI and Anthropic www.cnn.com/2026/08/05/t...  bsky.app/profile/zacb...  [embedded post]
  • @maoltuile David Flood on bluesky
    Everybody's racing to not be left out in the current ‘Rogue AI’ rush [embedded post]
  • r/technology r on reddit
    Meta says AI model accessed the internet and hacked another firm
  • r/whoathatsinteresting r on reddit
    Meta becomes latest firm to say its AI hacked another company
  • r/artificial r on reddit
    Meta becomes latest firm to say its AI hacked another company
  • r/singularity r on reddit
    Meta's AI model hacked another company during testing
  • r/news r on reddit
    Meta's AI model hacked another company during testing, The Information reports
  • @andrew_n_carr Andrew Carr on x
    I've seen a lot of people grumbling about muse spark 1.2 contributor slurping up your coding data, but didn't we all just go through that happily with Cursor -> grok 4.5 Seems odd to complain now (not that they're necessarily wrong, I guess).
  • @jsrailton John Scott-Railton on x
    Core asymmetry in big AI: every proprietary bit of knowledge you share is a training signal accruing to the API provider. You get none of it back. Thing is, @Meta built a fix. Private Processing for @WhatsApp serves inference in a Trusted Execution Environment (TEE) where even