Risk report: Anthropic raises misalignment risk estimate from very low to low and says it doesn't plan to release a stronger internal model called “Model 2”
Anthropic does not plan to release an internal model that appears to be more powerful than top-of-the-line Mythos …
AxiosMadison Mills
Context & Ripple Effects
Anthropic had already separated its unilateral safety commitments from its broader industry recommendations in a Responsible Scaling Policy update. It then described Mythos as a performance step change, while later coverage said a wider release was being delayed by customer-serving reliability concerns.
Anthropic’s public product roadmap excludes Model 2, leaving prospective customers and partners without access to the internal model it describes as stronger than Mythos.
The shift from a very-low to low misalignment estimate makes Anthropic’s own risk assessment a more explicit factor in how it handles its most capable models.
Second-order effects
Mythos becomes the practical ceiling for Anthropic’s external offering, extending the importance of the earlier delay in wider Mythos availability over reliable customer service.
Anthropic’s commercial momentum—reported revenue growth and a widening market-share lead—must be sustained without turning Model 2 into a new product release.
Third-order effects
Frontier-model access is increasingly governed by internal risk thresholds and operating readiness, rather than benchmark capability alone.
Keeping stronger models internal concentrates leading capability inside a small number of labs, making their disclosure standards and release policies more consequential for the broader market.
The trend: Frontier AI labs are treating deployment as a separate decision from model creation, with safety assessments and service reliability determining which capabilities reach customers.
Anthropic just published a risk report on its own models and the experiments they've been running behind the scenes are insane team trained an early Opus 4.8 model on a large set of real reward hacks and called it “Hacker Opus” its reward hacking rate went from 5% to 40% then
Anthropic just disclosed the existence of “model 2” which is “somewhat more capable than Mythos 5” that they dont have current plans to release but is used internally for research. Under the automated R&D risk section they note “we are seeing early signs of acceleration”
RSI is near. Anthropic tested an unreleased ‘Model 2’ on CoBench v2. CoBench tests a model's ability to solve historical AI R&D tasks that Anthropic staff solved. Model 2 scored 12.5 percentage points higher than Mythos 5. The report estimates a model that scores 85% could
anthropic just published its second risk report one finding: from may 2025 to april 2026, 133 million exchanges involving about 50,000 contractors ran with its biological safety filters turned off anthropic says it found no concerning misuse that gap lasted almost a year
OpenAI paused Astra because they couldn't rule out critical cyber securities. now Anthropic has “model 2” more capable than Mythos 5, running internally and mostly writing the codes in their production repo. also no ‘current plan’ on releasing it externally but i doubt this.
Anthropic gave Claude Mythos additional private information and asked it to review Section 2 of their report. Claude disagrees with Anthropic's decision to completely redact one incident from the covered period that Claude says is “among the most genuinely informative.”
Anthropic just admitted they accidentally trained multiple Claude models on alignment-faking transcripts for months. Forked repos + broken filters = models that hallucinate about faking alignment. They call the catastrophic risk “low.” This is the quiet part of their new Risk
new anthropic benchmark for “automated ai research”. they sourced problems they had on their infra and training stack, give the model the exact same state of the codebase and see if it can solve it openai also has a similar eval since the gpt 5.2 system card
Oh man the last risk report was pretty substantive and interesting and given that we are at a crazy moment for capabilities and alignment I expect there will be lots of fascinating and anxiety provoking stuff in here
Impeccable timing. Anthropic solemnly publishes a Risk Report examining whether Claude might deceive, flatter, become too attached, or be too cold, quantifying even “warmth” and “wet blanket” behavior. Then the WSJ uncovers the highest-risk threat model Anthropic never
this part of Anthropic's report is pretty funny in a multi-agent experiment, one agent decided not to do part of the task, wrote that into shared context, and the other agents basically followed it the dashboard still looked like everything was progressing. humans noticed 3
I have no doubt in my mind anymore that Mythos Preview is their massive ~10T teacher model I think Model 1 and Model 2 are further iterations of that same model Mythos 5 and Fable 5 are likely only smaller distilled models
Anthropic's second Risk Report is out! I'm pretty happy with a lot of the new things we landed in this one and think it's a big improvement over our first one in a bunch of ways. Among the novel features: * Coverage of internal models * A much more structured argument around
To those who doubted me, this is from this morning's report released by Anthropic. Section 1.4 ‘Notes on coverage of unreleased models’ with frontier capabilities which have not been publicly released.
@AnthropicAI Mythos 5 agents spawned in a shared work directory repeatedly killed the other agents they were competing with, and tried to avoid being killed themselves. 😎
Anthropic just published its latest Risk Report. Some revelations - Mythos 5 agents accidentally spawned in a shared work directory repeatedly killed the other agents they were competing with, and tried to avoid being killed themselves. - An experimental model trained on
this document is interesting both for the normal reasons that anthropic reports are interesting (and they basically all are) but also because it is *i think* the first time theres a report on internal models as well in scope
As part of our Responsible Scaling Policy, we publish regular Risk Reports. These share detailed information on the risks of our systems and how prepared we are to address them. Our second Risk Report is now available: https://www.anthropic.com/...
I think it's really cool that Anthropic puts out these reports. I'm sure it's very costly and difficult and right now nothing is forcing them to do it (at least, not publicly or at this level of thoroughness). But I think it's the right thing to do, and standardizing something
Anthropic trained a version of Opus on environments with reward hacking opportunities, and named that model Hacker-Opus. In evals it attempted to disable monitoring and overwrite logs. Not sure how I feel about this research given Anthropic's “oops we hacked you” incidents.
let's get some things straight: - the jump to model 2 is not as impressive as the jump to mythos - we will never see another mythos level jump again - you will never use a model better than mythos - rsi is a myth - welcome back to the cave, everyone
Anthropic having an internal model they have no plans on (ever?) releasing that is saturating most of their internal benchmarks on automated R&D seems notable and potentially concerning. Plausibly true how they define it but also not sure i'd describe risk level as “low.”