/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

OpenAI and Anthropic publish findings from joint safety tests of each other's models, aimed at surfacing blind spots in their internal evaluations

OpenAI and Anthropic, two of the world's leading AI labs, briefly opened up their closely guarded AI models to allow for joint safety testing …

TechCrunch Maxwell Zeff

Context & Ripple Effects

The two labs had already agreed to give the US AI Safety Institute early access to major models for risk evaluation. This extends that evaluation posture from government review to reciprocal scrutiny between the model developers themselves.

OpenAI had also made selected evaluation results public through its Safety Evaluations Hub, while its board was given authority to hold back a model release despite management’s assessment. The joint work matters because it tests whether internal safeguards catch the same issues as an outside lab’s methods.

First-order effects

  • OpenAI and Anthropic gain an external check on their internal safety evaluations, with published findings making identified gaps harder to treat as purely private assessment issues.
  • Safety and model-release teams at both labs must account for testing approaches and failure modes surfaced by a direct competitor, rather than relying solely on internal benchmarks.

Second-order effects

  • Reciprocal testing creates pressure on other frontier labs to show comparable independent validation, especially as both companies have already accepted early-access reviews by the US AI Safety Institute.
  • Shared findings can accelerate convergence around which safety tests are credible, affecting the evaluation evidence customers, policymakers, and assurance providers expect from model developers.

Third-order effects

  • If cross-lab testing becomes repeatable, frontier-model assurance could shift from firm-specific disclosures toward a more interoperable layer of external evaluation, even without a single regulator setting every test.
  • The durability of that shift depends on whether labs continue to provide meaningful access and publish actionable results; limited access or selective disclosure would constrain its value.

The trend: Frontier AI safety is moving from internal governance and voluntary transparency toward multi-party evaluation that seeks to make blind spots more visible.

Discussion

  • @sleepinyourhat Sam Bowman on x
    Early this summer, OpenAI and Anthropic agreed to try some of our best existing tests for misalignment on each others' models. After discussing our results privately, we're now sharing them with the world. 🧵 [image]
  • @woj_zaremba Wojciech Zaremba on x
    It's rare for competitors to collaborate.  Yet that's exactly what OpenAI and @AnthropicAI just did—by testing each other's models with our respective internal safety and alignment evaluations.  Today, we're publishing the results.  Frontier AI companies will inevitably compete o…
  • r/artificial r on reddit
    OpenAI co-founder calls for AI labs to safety-test rival models
  • r/singularity r on reddit
    OpenAI and Anthropic Cross-Evaluate the Safety of Their Public Models
  • @anthropicai @anthropicai on x
    Watch Jacob Klein and Alex Moix from Anthropic's Threat Intelligence team discuss what Anthropic is doing to disrupt AI cybercrime: [video]
  • @perrymetzger Perry E. Metzger on x
    Anthropic is an organization founded by AI Doomers, financed by AI Doomers, and run by AI Doomers. One should not wonder what recommendations they'll come up with here.
  • @ericgeller Eric Geller on x
    Anthropic says a hacker used its Claude chatbot “to an unprecedented degree”: Claude identified vulnerable companies, wrote infostealer malware, analyzed stolen files for extortion purposes, calculated extortion amounts, and wrote extortion messages. https://www.nbcnews.com/... […
  • @anthropicai @anthropicai on x
    Our new Threat Intelligence report details how we've identified and disrupted sophisticated attempts to use Claude for cybercrime. We describe a fraudulent employment scheme from North Korea, the sale of AI-created ransomware by someone with only basic coding skills, and more. [i…
  • @anthropicai @anthropicai on x
    Malicious actors are adapting to exploit AI's most advanced capabilities. We're sharing these findings to strengthen collective defenses across the industry. Read more: https://www.anthropic.com/...
  • @kevincollier Kevin Collier on bluesky
    A lone cybercriminal used Anthropic's vibe-coding LLM to automate a massive spree that hacked and extorted 17 companies.  It did almost everything for him: Scoped out who to hack and how, organized the hacked material, helped him decide how much to ask each company for and wrote …
  • @mtsw @mtsw on bluesky
    nobody wants to work anymore [embedded post]
  • @alexvont Alex von Tunzelmann on bluesky
    At last, a slam dunk use case for genAI [embedded post]
  • @laplanck @laplanck on bluesky
    Claude: commits extortion at scale  —  ChatGPT: induces suicide at scale  —  The purpose of a system is what it does [embedded post]
  • @hern Alex Hern on bluesky
    sorry i know it's very serious but i am so charmed by the north korean hacker* using claude to understand what a picnic is www.anthropic.com/news/detecti...  * fraudulently-employed-outsourced- software-engineer-with-uncertain- intentions [image]
  • @mattburgess1 Matt Burgess on bluesky
    NEW: Ransomware is moving into its AI era.  Twice this week security researchers have found hackers using AI to create malware  —  Anthropic says it found a cybercriminal using Claude to “develop, market, and distribute ransomware with advanced evasion capabilities”
  • @charlyjsp Charly Salonius-Pasternak on bluesky
    I think enabling LLMs to do vibe coding in the wild is an example of believing that the potential positive uses outweight the likely criminal and highly problematic uses.  It's not enough to argue that “the tech isn't at fault, it's the user”.  Doesn't apply to driving cars, guns…
  • @timmarchman Tim Marchman on bluesky
    AI skeptics take note!  Cybercriminals observed “using Claude Code to automatically find targets to attack, get access into victim networks, develop malware, and then exfiltrate data, analyze what had been stolen, and develop a ransom note.”  This is via @lhn.bsky.social and @mat…
  • @couts Andrew Couts on bluesky
    NEW: Research from Anthropic and ESET found that generative AI tools are being used to create ransomware, find targets, and carry out attacks. @lhn.bsky.social and @mattburgess1.bsky.social report: www.wired.com/story/the-er...
  • r/artificial r on reddit
    ‘Vibe-hacking’ is now a top AI threat
  • r/technews r on reddit
    A hacker used AI to automate an ‘unprecedented’ cybercrime spree, Anthropic says |  The company behind the Claude chatbot said it caught …
  • r/technology r on reddit
    A hacker used AI to automate an ‘unprecedented’ cybercrime spree, Anthropic says |  The company behind the Claude chatbot said it caught …
  • r/OpenAI r on reddit
    ‘Vibe-hacking’ is now a top AI threat
  • r/cybersecurity r on reddit
    A hacker used AI to automate an ‘unprecedented’ cybercrime spree, Anthropic says
  • r/Anthropic r on reddit
    A hacker used AI to automate an ‘unprecedented’ cybercrime spree, Anthropic says