/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

State of AI safety: as capabilities grow and models can monitor other models, issues like adversarial robustness persist and society is still not ready for AI

Here is a quick overview of my intuitions on where we are with AI safety in early 2026:  — So far, we continue to see exponential improvements in capabilities.

Windows On Theory Boaz Barak

Context & Ripple Effects

The safety debate has moved from baseline guardrails toward whether evaluation and oversight can keep pace with capability gains. Earlier coverage described evaluation and safeguard practices shaped by real-world use, while an ex-OpenAI researcher warned that competitive pressure can weaken alignment discipline as labs race toward more capable systems.

This assessment adds a practical tension: models may increasingly assist in monitoring other models, yet adversarial robustness remains unresolved and broader social readiness is still in question. That makes model-based oversight a complement to, rather than a replacement for, governance and testing.

First-order effects

  • AI safety teams gain a potentially more capable tool for reviewing model behavior, but must still treat model-generated monitoring as fallible rather than decisive.
  • Persistent adversarial-robustness problems keep deployment risk centered on how models behave under manipulation, even as capabilities improve.

Second-order effects

  • Labs deploying model-on-model monitoring will face pressure to demonstrate that the monitors are reliable against the same kinds of adversarial failures they are meant to detect.
  • The gap between technical monitoring capacity and societal preparedness raises the value of operational assurance processes, including evaluation, escalation, and accountable release decisions.

Third-order effects

  • If capability growth continues to outpace robustness and institutional readiness, AI safety is likely to shift from one-time pre-release checks toward continuous, layered oversight of deployed systems.
  • Model-assisted supervision could become standard infrastructure, but its credibility will depend on independent evaluation and governance rather than on automated monitoring alone.

The trend: AI safety is moving toward operational, model-assisted assurance while robustness and governance remain the limiting constraints on deployment confidence.

Discussion

  • @so8res Nate Soares on x
    Safety folks at the AI companies apparenly can't tell the difference between “the AI superficially does mostly what I ask” and the deep alignment properties that'd be needed for superintelligence, which casts doubt on their ability to pull off alignment.
  • @_nathancalvin Nathan Calvin on x
    Appreciate Sam endorsing this post which contains some pretty frank talk about the good bad and ugly of AI safety in 2026. I will keep saying that the actions of OpenAI's Global Affairs team (and related Super Pacs) do not seem consistent with taking these concerns seriously! [im…
  • @sama Sam Altman on x
    This is a very good post:
  • @_nathancalvin Nathan Calvin on x
    My views are similar. Alignment progress better than I expected (though still need lots more work, and better assurances that progress will remain robust). Societal readiness worse than I hoped. (Yet another 100m anti guardrails AI superpac announced on Sunday unlikely to help) […
  • @fleetingbits @fleetingbits on x
    @boazbaraktcs i think it is a mistake to think that there ever can be societal readiness for a disruptive technology before the disruptive effects are felt. governments can move very fast in a short time when faced with an obvious effect (e.g. 2008, covid) but not otherwise.
  • @boazbaraktcs Boaz Barak on x
    New blog post: the state of AI safety in four fake graphs. [image]