/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

State of AI safety: as capabilities grow and models can monitor other models, issues like adversarial robustness persist and society is still not ready for AI

Windows On Theory Boaz Barak

Context & Ripple Effects

The coverage arc has moved from public warnings about the risks of a high-profile conversational AI experiment to concerns that competitive pressure can erode alignment work. This assessment adds a technical constraint: model-on-model monitoring may improve oversight capacity, but it does not resolve adversarial robustness.

It also follows earlier evidence that AI safety governance is contested inside labs: OpenAI's board retained authority to block a release despite management's safety view. The enduring issue is whether governance and evaluation can keep pace with expanding capability.

First-order effects

  • AI developers and deployers must treat monitoring models as an additional safety layer, not as proof that underlying models are robust against adversarial inputs.
  • Unresolved robustness failures keep the burden on safety teams and release-governance bodies to test systems under hostile or unexpected conditions before broader use.

Second-order effects

  • Labs competing on capability face stronger pressure to demonstrate that their oversight methods work against manipulation, rather than simply adding automated monitors to deployment workflows.
  • Organizations considering consequential AI uses will need operational controls around monitoring outputs, since a monitor can itself be limited or misled by the system it is meant to oversee.

Third-order effects

  • If capabilities continue to outpace robust evaluation, AI governance is likely to shift from voluntary safety claims toward more operationally auditable testing and release controls.
  • The longer-run fault line is whether scalable oversight techniques can materially reduce failure risk; if not, pressure for slower deployment and stronger external accountability will grow.

The trend: This is one data point in the shift from aspirational AI-safety principles toward the practical problem of governing increasingly capable systems whose safeguards must withstand adversarial behavior.

Discussion

  • @_nathancalvin Nathan Calvin on x
    Appreciate Sam endorsing this post which contains some pretty frank talk about the good bad and ugly of AI safety in 2026. I will keep saying that the actions of OpenAI's Global Affairs team (and related Super Pacs) do not seem consistent with taking these concerns seriously! [im…
  • @so8res Nate Soares on x
    Safety folks at the AI companies apparenly can't tell the difference between “the AI superficially does mostly what I ask” and the deep alignment properties that'd be needed for superintelligence, which casts doubt on their ability to pull off alignment.
  • @sama Sam Altman on x
    This is a very good post:
  • @_nathancalvin Nathan Calvin on x
    My views are similar. Alignment progress better than I expected (though still need lots more work, and better assurances that progress will remain robust). Societal readiness worse than I hoped. (Yet another 100m anti guardrails AI superpac announced on Sunday unlikely to help) […
  • @fleetingbits @fleetingbits on x
    @boazbaraktcs i think it is a mistake to think that there ever can be societal readiness for a disruptive technology before the disruptive effects are felt. governments can move very fast in a short time when faced with an obvious effect (e.g. 2008, covid) but not otherwise.
  • @boazbaraktcs Boaz Barak on x
    New blog post: the state of AI safety in four fake graphs. [image]