OpenAI and Anthropic publish findings from joint safety tests of each other's models, aimed at surfacing blind spots in their internal evaluations
OpenAI and Anthropic, two of the world's leading AI labs, briefly opened up their closely guarded AI models to allow for joint safety testing …
TechCrunch Maxwell Zeff
Context & Ripple Effects
The two labs had already agreed to give the US AI Safety Institute early access to major models for risk evaluation. This extends that evaluation posture from government review to reciprocal scrutiny between the model developers themselves.
OpenAI had also made selected evaluation results public through its Safety Evaluations Hub, while its board was given authority to hold back a model release despite management’s assessment. The joint work matters because it tests whether internal safeguards catch the same issues as an outside lab’s methods.
First-order effects
- OpenAI and Anthropic gain an external check on their internal safety evaluations, with published findings making identified gaps harder to treat as purely private assessment issues.
- Safety and model-release teams at both labs must account for testing approaches and failure modes surfaced by a direct competitor, rather than relying solely on internal benchmarks.
Second-order effects
- Reciprocal testing creates pressure on other frontier labs to show comparable independent validation, especially as both companies have already accepted early-access reviews by the US AI Safety Institute.
- Shared findings can accelerate convergence around which safety tests are credible, affecting the evaluation evidence customers, policymakers, and assurance providers expect from model developers.
Third-order effects
- If cross-lab testing becomes repeatable, frontier-model assurance could shift from firm-specific disclosures toward a more interoperable layer of external evaluation, even without a single regulator setting every test.
- The durability of that shift depends on whether labs continue to provide meaningful access and publish actionable results; limited access or selective disclosure would constrain its value.
The trend: Frontier AI safety is moving from internal governance and voluntary transparency toward multi-party evaluation that seeks to make blind spots more visible.
Related: Operational AI assurance · The state-compatible AI lab · AI labs · US AI Safety Institute early access agreement · OpenAI Safety Evaluations Hub · OpenAI board model-release safeguards
Related Coverage
- Findings from a pilot Anthropic-OpenAI alignment evaluation exercise: OpenAI Safety Tests OpenAI
- Findings from a Pilot Anthropic—OpenAI Alignment Evaluation Exercise Anthropic
- The AI Wars: Data, Developers and the Battle for Market Share The New Stack · Pete Johnson
- Competitors OpenAI and Anthropic evaluated each other's AI security - which revealed shortcomings and advantages AIN · Vira Oliinyk
- OpenAI and Anthropic conducted safety evaluations of each other's AI systems Engadget · Anna Washenko
- OpenAI, Anthropic Joint Safety Tests Reveal Alarming Flaws in Rival AI Models WinBuzzer · Markus Kasanmascheff
- Teen Death Lawsuit: OpenAI Says ChatGPT's Safety Guardrails May Weaken in Longer Conversations Tech Times · Jose Enrico
- Anthropic and OpenAI Evaluate Safety of Each Other's AI Models PYMNTS.com
- OpenAI, Anthropic Team Up for Research on Hallucinations, Jailbreaking Bloomberg · Rachel Metz
- Threat Intelligence Report: August 2025 Anthropic
- AI firm says its technology weaponised by hackers BBC · Imran Rahman-Jones
- A hacker used AI to automate an ‘unprecedented’ cybercrime spree, Anthropic says NBC News · Kevin Collier
- Detecting and countering misuse of AI: August 2025 Anthropic
- ‘Vibe hacking’ is here as Anthropic reveals as Claude AI being used in extortion and ransomware schemes Moneycontrol
- The Era of AI-Generated Ransomware Has Arrived Wired
- Researchers uncover first AI-generated ransomware capable of writing its own attack code GreenBot · Michael Anthony Bitoon
- Hackers Attempted to Misuse Claude AI to Launch Cyber Attacks Cyber Security News · Mayura Kathir
- Agentic AI coding assistant helped attacker breach, extort 17 distinct organizations Help Net Security · Zeljka Zorz
- Hacker Used AI To Launch ‘Unprecedented’ Cyberattack — and It Could Happen Again Tom's Guide · Amanda Caswell
- Vibe-hacking based AI attack turned Claude against its safeguard: Here's how Digit · Vyom Ramani
- Start Up No.2503: Anthropic's Claude helps hacker's extortion, can Democrat influencers.. influence?, VR retail woes, and more The Overspill · Charlesarthur
- Why AI Keeps Security Experts Awake At Night? NDTV Profit · Ivor Soans
- Anthropic says its chatbot Claude used in extortion, fraud schemes Daily Sabah
- A hacker turned a popular AI tool into a cybercrime machine Digital Trends · Trevor Mogg
- Anthropic Says ‘Vulnerabilities Need Fixing’ for Claude for Chrome Before Public Launch Analytics India Magazine · Supreeth Koundinya
- Anthropic thwarts hacker attempts to misuse Claude AI for cybercrime Reuters
- Anthropic says agentic AI ran an extortion playbook, end to end Implicator.ai · Maria Garcia
- Criminals are ‘vibe hacking’ with AI at unprecedented levels: Anthropic Cointelegraph · Brayden Lindrea
- Anthropic Says Bad Actors Have Now Turned To ‘Vibe Hacking’ BGR · Joshua Hawkins
- Mystery Hacker Used AI To Automate ‘Unprecedented’ Cybercrime Rampage ZeroHedge News · Tyler Durden
- Crims laud Claude to plant ransomware and fake IT expertise The Register · Thomas Claburn
- Anthropic Stops Hacker From Using Claude AI to Breach Companies Tech.co · Conor Cawley
- ‘Vibe Hacking’: Criminals Are Weaponizing AI With Help From Bitcoin, Says Anthropic Yahoo Finance · Josh Quittner
- ‘Vibe Hacking’: Criminals Are Weaponizing AI With Help From Bitcoin, Says Anthropic Decrypt · Josh Quittner
- Chatbot's Crime Spree Used AI to Grab Bank Details, Social Security Numbers Gizmodo · Riley Gutiérrez McDermid
- Anthropic admits its AI is being used to conduct cybercrime Engadget · Matt Tate
- Anthropic: Vibe hacking and weaponizing agentic AI Constellation Research · Larry Dignan
- Anthropic Says Attacker Used AI Tool in Widespread Hacks Bloomberg · Emily Forgash
- Anthropic Warns of New ‘Vibe Hacking’ Attacks That Use Claude AI CNET · Omar Gallaga
- Anthropic forms national security advisory council to guide AI use in government Reuters · Akash Sriram
- AI is becoming a core tool in cybercrime, Anthropic warns Help Net Security · Anamarija Pogorelec
- Anthropic says cybercriminals used its Claude AI for ‘vibe hacking’ Business Insider · Robert Scammell
- #Anthropic has come out and said that a criminal was able to target and extort numerous companies utilizing AI (#Claude in this case). … Al Pascual
- A recent Threat Intelligence Report (August 2025) from Anthropic revealed how fraudsters are embedding AI across every stage of their operations … Ang Sebastian
- Sharing this excellent threat intelligence report from Anthropic showing the rise in threat actors use of AI in their tactics. Thanks for the heads up Pedram Amini … Greg Martin
- Detecting and Countering Misuse of AI — Anthropic just shared their Threat Intelligence Report for August 2025, and it's a doozy. … Chris H.
- Great report from Anthropic out today regarding threat actor use of Claude Code. Not surprised to see that adversaries have adopted vibecoding & LLMs to speed up their operations. … Will Urbanski
- NBC News: A hacker used AI to automate an ‘unprecedented’ cybercrime spree, Anthropic says — The company behind the Claude chatbot said it caught a hacker using its chatbot to identify, hack and extort at least 17 companies. — https://www.nbcnews.com/... @briankrebs@infosec.exchange · BrianKrebs
- Anthropic has released a 25-page Threat Intelligence report that looks like really good reading. Very specific, verified case studies of what they're seeing in the wild. — “Detecting and countering misuse of AI: August 2025” — #infosec #cybersecurity #threatintel — https://www.anthropic.com/... @neurovagrant@masto.deoan.org · Ian Campbell
Discussion
-
@sleepinyourhat
Sam Bowman
on x
Early this summer, OpenAI and Anthropic agreed to try some of our best existing tests for misalignment on each others' models. After discussing our results privately, we're now sharing them with the world. 🧵 [image]
-
@woj_zaremba
Wojciech Zaremba
on x
It's rare for competitors to collaborate. Yet that's exactly what OpenAI and @AnthropicAI just did—by testing each other's models with our respective internal safety and alignment evaluations. Today, we're publishing the results. Frontier AI companies will inevitably compete o…
-
r/artificial
r
on reddit
OpenAI co-founder calls for AI labs to safety-test rival models
-
r/singularity
r
on reddit
OpenAI and Anthropic Cross-Evaluate the Safety of Their Public Models
-
@anthropicai
@anthropicai
on x
Watch Jacob Klein and Alex Moix from Anthropic's Threat Intelligence team discuss what Anthropic is doing to disrupt AI cybercrime: [video]
-
@perrymetzger
Perry E. Metzger
on x
Anthropic is an organization founded by AI Doomers, financed by AI Doomers, and run by AI Doomers. One should not wonder what recommendations they'll come up with here.
-
@ericgeller
Eric Geller
on x
Anthropic says a hacker used its Claude chatbot “to an unprecedented degree”: Claude identified vulnerable companies, wrote infostealer malware, analyzed stolen files for extortion purposes, calculated extortion amounts, and wrote extortion messages. https://www.nbcnews.com/... […
-
@anthropicai
@anthropicai
on x
Our new Threat Intelligence report details how we've identified and disrupted sophisticated attempts to use Claude for cybercrime. We describe a fraudulent employment scheme from North Korea, the sale of AI-created ransomware by someone with only basic coding skills, and more. [i…
-
@anthropicai
@anthropicai
on x
Malicious actors are adapting to exploit AI's most advanced capabilities. We're sharing these findings to strengthen collective defenses across the industry. Read more: https://www.anthropic.com/...
-
@kevincollier
Kevin Collier
on bluesky
A lone cybercriminal used Anthropic's vibe-coding LLM to automate a massive spree that hacked and extorted 17 companies. It did almost everything for him: Scoped out who to hack and how, organized the hacked material, helped him decide how much to ask each company for and wrote …
-
@mtsw
@mtsw
on bluesky
nobody wants to work anymore [embedded post]
-
@alexvont
Alex von Tunzelmann
on bluesky
At last, a slam dunk use case for genAI [embedded post]
-
@laplanck
@laplanck
on bluesky
Claude: commits extortion at scale — ChatGPT: induces suicide at scale — The purpose of a system is what it does [embedded post]
-
@hern
Alex Hern
on bluesky
sorry i know it's very serious but i am so charmed by the north korean hacker* using claude to understand what a picnic is www.anthropic.com/news/detecti... * fraudulently-employed-outsourced- software-engineer-with-uncertain- intentions [image]
-
@mattburgess1
Matt Burgess
on bluesky
NEW: Ransomware is moving into its AI era. Twice this week security researchers have found hackers using AI to create malware — Anthropic says it found a cybercriminal using Claude to “develop, market, and distribute ransomware with advanced evasion capabilities”
-
@charlyjsp
Charly Salonius-Pasternak
on bluesky
I think enabling LLMs to do vibe coding in the wild is an example of believing that the potential positive uses outweight the likely criminal and highly problematic uses. It's not enough to argue that “the tech isn't at fault, it's the user”. Doesn't apply to driving cars, guns…
-
@timmarchman
Tim Marchman
on bluesky
AI skeptics take note! Cybercriminals observed “using Claude Code to automatically find targets to attack, get access into victim networks, develop malware, and then exfiltrate data, analyze what had been stolen, and develop a ransom note.” This is via @lhn.bsky.social and @mat…
-
@couts
Andrew Couts
on bluesky
NEW: Research from Anthropic and ESET found that generative AI tools are being used to create ransomware, find targets, and carry out attacks. @lhn.bsky.social and @mattburgess1.bsky.social report: www.wired.com/story/the-er...
-
r/artificial
r
on reddit
‘Vibe-hacking’ is now a top AI threat
-
r/technews
r
on reddit
A hacker used AI to automate an ‘unprecedented’ cybercrime spree, Anthropic says | The company behind the Claude chatbot said it caught …
-
r/technology
r
on reddit
A hacker used AI to automate an ‘unprecedented’ cybercrime spree, Anthropic says | The company behind the Claude chatbot said it caught …
-
r/OpenAI
r
on reddit
‘Vibe-hacking’ is now a top AI threat
-
r/cybersecurity
r
on reddit
A hacker used AI to automate an ‘unprecedented’ cybercrime spree, Anthropic says
-
r/Anthropic
r
on reddit
A hacker used AI to automate an ‘unprecedented’ cybercrime spree, Anthropic says