State of AI safety: as capabilities grow and models can monitor other models, issues like adversarial robustness persist and society is still not ready for AI
Here is a quick overview of my intuitions on where we are with AI safety in early 2026: — So far, we continue to see exponential improvements in capabilities.
Context & Ripple Effects
The safety debate has moved from baseline guardrails toward whether evaluation and oversight can keep pace with capability gains. Earlier coverage described evaluation and safeguard practices shaped by real-world use, while an ex-OpenAI researcher warned that competitive pressure can weaken alignment discipline as labs race toward more capable systems.
This assessment adds a practical tension: models may increasingly assist in monitoring other models, yet adversarial robustness remains unresolved and broader social readiness is still in question. That makes model-based oversight a complement to, rather than a replacement for, governance and testing.
First-order effects
- AI safety teams gain a potentially more capable tool for reviewing model behavior, but must still treat model-generated monitoring as fallible rather than decisive.
- Persistent adversarial-robustness problems keep deployment risk centered on how models behave under manipulation, even as capabilities improve.
Second-order effects
- Labs deploying model-on-model monitoring will face pressure to demonstrate that the monitors are reliable against the same kinds of adversarial failures they are meant to detect.
- The gap between technical monitoring capacity and societal preparedness raises the value of operational assurance processes, including evaluation, escalation, and accountable release decisions.
Third-order effects
- If capability growth continues to outpace robustness and institutional readiness, AI safety is likely to shift from one-time pre-release checks toward continuous, layered oversight of deployed systems.
- Model-assisted supervision could become standard infrastructure, but its credibility will depend on independent evaluation and governance rather than on automated monitoring alone.
The trend: AI safety is moving toward operational, model-assisted assurance while robustness and governance remain the limiting constraints on deployment confidence.