OpenAI says GPT-5.6 Sol and Terra were capable of identifying vulnerabilities but were unable to execute autonomous, end-to-end attacks against hardened targets
GPT-5.6 is a new family of three models: Sol, our new flagship model; Terra, a capable lower-cost option; and Luna, our fastest and most cost-efficient model.
Context & Ripple Effects
This assessment arrives alongside OpenAI’s staged introduction of the GPT-5.6 family: Sol as the flagship, Terra as a lower-cost option, and Luna as the speed- and cost-focused model. Related coverage describes an initial limited preview to roughly 20 companies with U.S. government disclosure, followed by plans for a broader rollout.
The surrounding coverage also emphasizes Sol’s ExploitBench performance and an Ultra mode using subagents for complex workflows. The security finding therefore qualifies the practical meaning of those capabilities: vulnerability discovery has advanced, while autonomous compromise of hardened targets was not demonstrated.
First-order effects
- OpenAI can position Sol and Terra as models that can assist with finding vulnerabilities without claiming they can independently carry out complete attacks on hardened systems.
- Organizations evaluating the preview receive a more specific risk boundary for these models: security-relevant analysis is in scope, but autonomous end-to-end offensive execution remains constrained in OpenAI’s reported testing.
Second-order effects
- Prospective customers and internal security teams are likely to separate vulnerability-identification use cases from autonomous remediation or offensive-security workflows when setting access controls and evaluation criteria.
- The combination of strong benchmark performance, subagent workflows, and a stated limit on autonomous attacks raises the importance of testing model behavior in realistic hardened environments rather than relying on capability benchmarks alone.
Third-order effects
- If model releases increasingly pair stronger agentic reasoning with explicit cyber-risk evaluations, security assessments may become a more central gate for enterprise deployment and public rollout decisions.
- The durable industry shift is toward treating vulnerability discovery and autonomous exploitation as distinct capability thresholds; whether that separation holds will depend on continued testing as models and agent tooling improve.
The trend: Frontier-model deployment is moving toward capability-specific cyber-risk disclosure as vendors commercialize increasingly agentic systems.