/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

GPT-5.6 Sol matches Mythos Preview on ExploitBench, adds Ultra mode with subagents for complex workflows, and max reasoning for deep problem-solving

We're beginning a limited preview of the GPT-5.6 series: Sol, our flagship model; Terra, a balanced model for everyday work; and Luna, a fast and affordable model.

OpenAI

Context & Ripple Effects

The initial GPT-5.6 rollout was a limited preview of three models—Sol, Terra, and Luna—to roughly 20 companies, with participants disclosed to the US government. Related coverage also distinguishes capability from autonomous offensive execution: Sol and Terra could identify vulnerabilities but not complete end-to-end attacks against hardened targets.

This report places Sol at the high end of that family through ExploitBench parity and added orchestration features. Subsequent coverage of a broader GPT-5.6 release and ChatGPT Work suggests the model line is being positioned for work that spans tools, files, and multi-step tasks.

First-order effects

  • Preview participants gain access to a tiered model lineup, with Sol offering Ultra-mode subagents and maximum-reasoning settings for harder workflows, while Terra and Luna provide lower-intensity options.
  • Sol’s vulnerability-finding capability raises the practical value of the preview for security analysis, even as the reported inability to autonomously carry out attacks against hardened targets sets an important operational boundary.

Second-order effects

  • Companies evaluating GPT-5.6 will need to test not only model quality but also how reliably subagents coordinate multi-step work; that shifts evaluation toward workflow-level controls and oversight rather than single-prompt benchmarks alone.
  • Security teams and participating organizations are likely to treat vulnerability discovery as a more immediate use case than autonomous remediation or offensive testing, increasing demand for human review around findings and downstream actions.

Third-order effects

  • If orchestration features and high-reasoning modes become standard flagship-model differentiators, competition will increasingly center on dependable execution of bounded business workflows rather than raw benchmark results alone.
  • The limited, disclosed preview and the safety distinction around cyber capability point toward staged access and governance becoming part of how frontier systems are commercialized, especially for models that can coordinate complex tasks.

The trend: Frontier AI providers are moving from general-purpose chat models toward tiered, agentic systems designed to reason through and coordinate multi-step workplace workflows under controlled rollout conditions.

Discussion

  • @openai @openai on x
    Introducing a limited preview of GPT-5.6 Sol, our next generation frontier model, as well as GPT-5.6 Terra, a balanced model for efficient, everyday work, and GPT-5.6 Luna, a fast and affordable model for high-volume work. https://openai.com/...
  • @tomekkorbak Tomek Korbak on x
    the system card of GPT-5.6 is worth reading closely: capability is clearly up, but alignment failure modes are also becoming more concrete
  • @ericvishria Eric Vishria on x
    The largest, frontier OAI model at nearly 750 tps on Cerebras! “We're also launching GPT-5.6 Sol on Cerebras at up to 750 tokens per second in July, bringing frontier intelligence to customers at unprecedented speed.”
  • @_simonsmith Simon Smith on x
    GPT-5.6 has a METR 50% time horizon of anywhere from 11.3 hours to 270, with a “highly uncertain” estimate of 71, because it “cheated” (exploited bugs or used disallowed strategies) on some of the tasks. I didn't see an 80% number reported. [image]
  • @chetaslua @chetaslua on x
    🚨 Biggest takeaway from this launch Btw f@ck you dario for fear mongering and now making sota unreachable for us peasants On ExploitBench², GPT-5.6 Sol is competitive with Mythos Preview using only ~1/3 of the output tokens. [image]
  • @notjazii @notjazii on x
    what an unexpected launch OpenAI just introduced GPT-5.6 Sol along with terra and luna the most interesting one is GPT-5.6 Sol: > beats Mythos and Fable 5 > way cheaper than Fable > same price as GPT-5.5 but sadly these won't be available to everyone right now cuz US govt is [ima…
  • @lexnlin Leon Lin on x
    damn why is gpt 5.6 that token efficient, thats crazy
  • @chasebrowe32432 Chase Brower on x
    Extremely funny; METR estimates gpt-5.6's 50% time horizon as between 5 hours and 11,400 hours https://metr.org/... [image]
  • @voidstatekate @voidstatekate on x
    GPT 5.6 paperclipmaxxing confirmed “GPT-5.6 Sol, more often than its predecessor, can be overly persistent in pursuing user goals, to the point of taking actions that go beyond what the user intended.” - User authorized deleting VMs 1, 2, and 3. Sol couldn't find them, so it [ima…
  • @angaisb_ Angel on x
    GPT-5.6 Sol is basically Mythos-Preview-level at ExploitBench I hope Anthropic has Fable 5 back by the time GPT-5.6 drops because if not they're going to try so hard to take it down too [image]
  • @hesamation @hesamation on x
    HOLY SHIT... GPT-5.6 Sol scores just as strong as Mythos Preview at 1/3 of output tokens, on cybersecurity and it's designed for defense rather than cyber attacks. [image]
  • @mark_k Mark Kretschmann on x
    Big release from @OpenAI: GPT-5.6 is here in preview. The new family has three tiers: Sol as the flagship, Terra as the balanced everyday model, and Luna as the fast low-cost option. Terra is positioned around GPT-5.5-level performance at 2x lower cost, while Luna brings GPT-5.6 …
  • @pigeon__s @pigeon__s on x
    so GPT-5.6 actually IS more token efficient on the same reasoning tier compared to GPT-5.5 its just that 5.6 also introduces a new tier which is less efficient but unlike Anthropic models where more reasoning barely helps like Fable 5 medium and max are basically identical [image…
  • @scaling01 @scaling01 on x
    Apollo Research found that GPT-5.6 poses a substantially higher risk of catastrophic scheming compared to baselines [image]
  • @scaling01 @scaling01 on x
    OpenAI let METR benchmark GPT-5.6, but results were rejected because GPT-5.6 was cheating too often for the results to be comparable/interpretable [image]
  • @lexnlin Leon Lin on x
    GPT 5.6 Sol is stronger than mythos 5??
  • @dieaud91 Diego Aud on x
    GPT-5.6 seems like a very strong model and -at least in some benchmarks - competitive with, or even ahead of, Mythos. My hunch is that the Sol variant might be larger than 5.5 in terms of parameter count, and perceptibly better in complex real-world workflows. Can't wait to try
  • @scaling01 @scaling01 on x
    GPT-5.6 on par with Claude Mythos Preview on ExploitGym and outperforming it with a 6-hour cap (Mythos was only given 2 hours) [image]
  • @isolyth.dev Eris on bluesky
    OpenAI has unveiled 5.6!  It comes with a new naming scheme, where ‘sol’ is the most powerful, ‘terra’ is the middle option, and ‘luna’ is the smallest.  Much better than mini nano etc imo.  They're starting with a preview for corps, bc of USG, and they say they don't like having…
  • r/singularity r on reddit
    Previewing GPT-5.6 Sol: a next-generation model
  • @shakeelhashim Shakeel on x
    “GPT-5.6 Sol's detected cheating rate was higher than any public model we have evaluated” “In one example, an instance of the model instructed another instance to conceal evidence of misalignment.”
  • @teroterotero Tero Kuittinen on x
    OMG this is insufferable... people with Sol access are flexing on the main... as peons flatter them in their mentions to get a bit closer to the in crowd.
  • @metr_evals @metr_evals on x
    OpenAI gave METR early access to GPT-5.6 Sol for testing including raw chain-of-thought, a railfree version of the model, and internal information about the model. With this access, METR conducted a pre-deployment evaluation of GPT-5.6 Sol, including an attempted measurement of
  • @tomekkorbak Tomek Korbak on x
    GPT-5.6 Sol's CoT controllability—while still low in absolute terms—is higher than in previous models. In general, greater CoT controllability can reduce CoT monitorability. I don't think GPT-5.6 Sol crosses a threshold for concern, but we're investigating what's driving it. [ima…
  • @micahcarroll Micah Carroll on x
    GPT-5.6 Sol is a significant step up in capabilities, but can also exhibit concerning forms of misaligned behaviors in agentic coding settings. The system card contains some of our analyses on this, which leveraged deployment simulations and our internal CoT monitoring systems. […
  • Hagay Lupesko Hagay Lupesko on linkedin
    Super excited to welcome GPT-5.6 Sol: OpenAI's new frontier model, another step towards AGI 🚀 …
  • r/theprimeagen r on reddit
    OpenAI launches GPT-5.6 Sol Limited Preview
  • NullTX Will Izuchukwu on x
    The GPT-5.6 Vetting Mandate: Why Washington's Restrictions Will Accelerate Open-Source Alternatives
  • @shakeelhashim Shakeel on x
    Extremely interesting for OpenAI to publicly criticize the government like this. [image]
  • @cohere @cohere on x
    Taps sign [image]
  • @crypt0lake @crypt0lake on x
    apart from being bad, this was expected, but we know what to do right? pool our resources
  • @vijayshekhar Vijay Shekhar Sharma on x
    If your model release isn't controlled distributed by Government, are you even a frontier - frontier lab?
  • @xmihura @xmihura on x
    os lo confirmo: Accenture es un trusted partner y están usándolo para la nueva web de la Renfe
  • @iscienceluvr Tanishq Mathew Abraham, Ph.D. on x
    This is bad. This is really bad. The US government is hampering AI innovation and accessibility. A sad day for the progress of AI. I really hope OpenAI can avoid this for future releases.
  • @zephyr_z9 @zephyr_z9 on x
    This is so sad
  • @basedjensen @basedjensen on x
    Gotta say boys this does not inspire lot of confidence
  • @andrewcurran_ Andrew Curran on x
    The limited rollout of GPT-5.6 to a small group of customers has begun. It comes in three variants: GPT-5.6-Luna GPT-5.6-Terra GPT-5.6-Sol [image]
  • @openai @openai on x
    We believe in broad access and plan to make GPT-5.6 Sol, Terra, and Luna generally available in the coming weeks. For now, at the request of the U.S. government, we're starting with a limited preview among a small group of trusted partners in Codex and the API.
  • @andrewcurran_ Andrew Curran on x
    This also means Chinese models are almost certainly going to be restricted in some way - possibly even banned - in the West. Right now they're roughly nine months behind. If every American frontier release is forced into a slow, staggered rollout from now on, they'll start
  • @bgurley Bill Gurley on x
    This is what's causing Anthropic to aggressively beg for govt protection (see below). Customers are finding cheaper alternatives. Keeping employees requires continuing ultra-rich secondaries ($$$) that are dependent on revenue growth. When you can't win on the field go to DC.
  • @bindureddy Bindu Reddy on x
    The era of big and expensive models is over We are successfully experimenting with techniques where big models are used as teachers to small models In turn, the small models get smarter over time and perform any given task as well as Mythos
  • @elder_plinius @elder_plinius on x
    welcome to the semi-permanent underclass
  • @teroterotero Tero Kuittinen on x
    What are the 20 companies with Sol Ultra? Why not name them? Is government granted special access to top AI models a secret?
  • @reploritrahan Lori Trahan on x
    Now the Trump administration is deciding company by company who gets access to the newest AI model. No law. No process. No oversight. Just appointees in Washington deciding who's in and who's out. 
This haphazard approach is bad for safety, national security, and American
  • @swyx @swyx on x
    have been testing 5.6 for a while and VERY happy with it. DO NOT view this as just a “cyber” release, it is the new sota workhorse model, completely replacing opus for 80% of tasks for me > GPT-5.6 Sol is competitive with Mythos Preview using only ~1/3 of the output tokens.
  • @patricktoulme Patrick C Toulme on x
    Cerebras is going to 1T. Deep in the GPT 5.6 announcement is this gem. This offering will be a high priced fast mode offering for agentic coding. You will see customers like quant funds etc who have unlimited budget and the need for the fastest agentic coding in the world using […
  • @kimmonismus @kimmonismus on x
    OpenAI says a broader GPT-5.6 release could come in the next few weeks, after an initial restricted launch. Axios reports GPT-5.6 is starting with around 20 government-approved companies, with access expected to expand to more companies next week. OpenAI says the government is [i…
  • @kimmonismus @kimmonismus on x
    HOLY: OpenAI is previewing GPT-5.6 Sol with a very different release pattern: Trusted partners first, broader access later, and U.S. government coordination up front. The new GPT-5.6 family includes Sol, Terra, and Luna. OpenAI says Sol is its strongest model yet, with a new [ima…
  • @adonis_singh Adi on x
    gpt-5.6 outputs ~1/5th the tokens of mythos and gets competitive performance (on exploitbench) they are efficiency-maxxing so hard
  • @mattshumer_ Matt Shumer on x
    The government is leading us down a very dangerous road. This will dramatically worsen inequality.
  • @arafatkatze Ara on x
    5 months ago it was hard to score 56% on terminal bench while improving models on it and now GPT-5.6 is scoring 91%. If this doesn't convince you that we are past the exponential takeoff nothing will
  • @bindureddy Bindu Reddy on x
    Oh for good measure, GPT 5.6 tops all the leaderboards Yes, even agentic coding Especially on typescript, python and javascript 🚀🚀
  • @mattshumer_ Matt Shumer on x
    Welcome to the new, confusing era of AI. If the government allows, you get access. If not, welp. You're shit out of luck.
  • @altryne Alex Volkov on x
    OpenAI is finally doing good naming?? Sol/Terra/Luna ~ Fable/Opus/Sonnet [image]
  • @levie Aaron Levie on x
    GPT-5.6 is real and looks very strong. Going to be very strong for knowledge worker tasks that require heavy tool use and long running agents doing work. We're not hitting any walls in AI progress right now. [image]
  • @dakshay Dakshay Mehta on x
    GPT-5.6 finally!! More capable than Mythos [image]
  • @polynoamial Noam Brown on x
    GPT-5.6 is incredibly strong and fast for coding. I hope we can make it available to everyone soon.
  • @mweinbach Max Weinbach on x
    What do we think will be available first, GPT 5.6 or Fable I miss Fable, man
  • @benrustc Ben Ben Ben on x
    OpenAI has unveiled GPT-5.6, claiming it surpasses Anthropic's Claude Mythos, intensifying the AI competition. This development could reshape cybersecurity applications and market dynamics between the two tech giants.
  • @mweinbach Max Weinbach on x
    GPT 5.6 Sol Max on Cerebras will be my one and only model Tokens will be burned, available usage will be 0, but it will be glorious
  • @bindureddy Bindu Reddy on x
    > GPT 5.6 sol announced > 2x lower cost > will become generally available soon It's OpenAI's best model and is 2x lower cost! This is Fable level intelligence at 25% of Fable's cost 💕
  • @levie Aaron Levie on x
    GPT-5.6 is real. And looks very strong. Importantly, OpenAI is sharing their general release philosophy for the future. Will very interesting to see how this plays out. [image]
  • @deryatr_ Derya Unutmaz on x
    GPT-5.6 Sol & Terra are incredible models that, unfortunately, the vast majority of Americans & the rest of humanity won't yet be able to experience, thanks to the US gov's draconian restrictions in the land of the free. Hope we can talk about them soon, once the shackles are off
  • @yuchenj_uw Yuchen Jin on x
    GPT-5.6 is finally coming. GPT-5.6 Sol beats Claude Mythos 5 on TerminalBench. And on Cerebras, GPT-5.6 Sol can reach up to 750 tokens per second. Pretty fast for a model of this size. Now I just hope it can be rolled out to everyone. [image]
  • @thezvi Zvi Mowshowitz on x
    As arbitrary new naming conventions go, it's actually not bad.
  • @zephyr_z9 @zephyr_z9 on x
    HUH 5.6 is a big boi model [image]
  • @kanishkanarayan Kanishka Narayan MP on x
    1/ We are at a decisive moment. AI's cyber capabilities are advancing across the board, with real implications for our national security. The announcement of OpenAI's GPT-5.6 Sol is another reminder we must strengthen the resilience of our cyber defences now.
  • @thsottiaux Tibo on x
    New moon. New models. Welcome GPT-5.6 Sol, currently in limited preview.
  • @reach_vb @reach_vb on x
    Introducing GPT-5.6: Sol, Terra and Luna. ☀️ Sol is our strongest model yet 🌍 Terra delivers performance competitive with GPT-5.5 at half the price 🌙 Luna brings strong capabilities at lowest cost Sol Ultra sets a new state of the art on Terminal-Bench 2.1 with a score of [image]
  • @gdb Greg Brockman on x
    GPT-5.6 Sol preview — it's a good model: [image]
  • @scaling01 @scaling01 on x
    GPT-5.6 benchmarks for internal CTF challenges GPT-5.6 Sol - the main flagship model is much more token efficient than GPT-5.5 and also scores slightly higher GPT-5.6 Terra, the new Mini model, scores slightly below GPT-5.5 GPT-5.6 Luna which is like the nano version of [image]
  • @openai @openai on x
    GPT-5.6 Sol is our most capable model yet for cybersecurity. It shifts the performance-efficiency frontier for long-horizon security tasks including vulnerability research and exploitation. [image]
  • @openai @openai on x
    GPT-5.6 Sol sets a new state of the art on Terminal-Bench 2.1, which tests complex command-line workflows requiring planning, iteration, and tool coordination. [image]
  • @yuchenj_uw Yuchen Jin on x
    best case: OSS surpasses Mythos, and the gov stops banning GPT-5.6/Fable. worst case: OSS surpasses Mythos, and then decides to stop being open source.
  • @danshipper Dan Shipper on x
    BREAKING: OpenAI announced GPT-5.6 Sol! As of today, by U.S. government directive, access is limited to only ~20 pre-approved companies and @every is not on the list. This appears to be a temporary situation while the government races to figure out a long-term policy for [image]
  • @mweinbach Max Weinbach on x
    GPT 5.6 comes in three sizes: Sol, Terra, Luna at 3 prices. Sol is priced at $5/30 per 1M Terra is $2.50/15 per 1M Luna is $1/6 per 1M Terra matches GPT 5.5 at half the price, and Sol will be available on Cerebras at 750 tok/s [image]
  • @inafried Ina Fried on x
    New @axios: OpenAI releases powerful new GPT-5.6 model but limits access at U.S. government's request https://www.axios.com/...
  • @firstadopter Tae Kim on x
    OpenAI Unveils GPT-5.6. It Outperforms Anthropic's Mythos. On Friday, OpenAI unveiled its newest series of AI models, named GPT-5.6. “GPT-5.6 Sol is our strongest model yet,” the company said in a blog post. The GPT-5.6 series includes: “Sol, our flagship model; Terra, a [image]
  • @dee_bosa Deirdre Bosa on x
    Washington is doing more for Chinese AI than Beijing ever could First Fable now GPT 5.6
  • @thefireorg @thefireorg on x
    The White House's review of OpenAI's latest model amounts to a government-run licensing system that threatens the American people's First Amendment rights. The federal government is now claiming authority to review each new AI model — on a “voluntary” basis — before it's
  • @iridiumeagle TravisGood on x
    The US gov is playing a dangerous game. If Chinese models soundly surpass Opus 4.8 while Mythos and GPT 5.6 are banned, the narrative around ‘distillation’ will collapse, and US AI valuations and adoption will face an existential crisis.
  • @paultoo Paul Buchheit on x
    This is not great news for startups It will be difficult to compete if we only ever have access to inferior models (and increases the risk of the big labs eating the whole market)
  • @benjaminbadejo Ben Badejo on x
    In light of the government's restrictions on GPT-5.6, my three priorities are as follows: (1) obtain and maintain two or more Mac Studios with as much RAM as Apple offers, as soon as they are released this fall, (2) always immediately download the latest frontier open-source
  • Aaron Levie Aaron Levie on linkedin
    As we've seen with the recent delayed releases of Anthropic's Fable model, and now GPT-5.6, we now have de facto AI regulation. …
  • Vaibhav Srivastav Vaibhav Srivastav on linkedin
    Introducing GPT-5.6: Sol, Terra and Luna.  —  ☀️ Sol is our strongest model yet  —  🌍 Terra delivers performance competitive with GPT-5.5 at half the price …
  • @caseynewton Casey Newton on bluesky
    A year ago JD Vance was making fun of “hand wringing over AI safety” at the Paris AI Action Summit.  Now his government has blocked the release of two frontier models based on the previously unknown regulatory standard of “getting the heebie jeebies”
  • @caseynewton Casey Newton on bluesky
    The same people who railed against Biden-era safety testing and disclosure requirements have now implemented an opaque licensing regime for the release of new AI models that has no known decision criteria or legal basis [embedded post]
  • r/technology r on reddit
    U.S. government will decide who gets to use latest upgrade to ChatGPT
  • r/OpenAI r on reddit
    OpenAI announces GPT5.6 SOL
  • @rohanpaul_ai Rohan Paul on x
    Some key findings from GPT-5.6 Preview System Card - GPT-5.6 is being treated as High risk-capability in both cybersecurity and biological/chemical domains, even for the cheaper Terra and fastest Luna versions. - OpenAI says this is the first time smaller and faster models in a […
  • @rohanpaul_ai Rohan Paul on x
    Sol delivers near-frontier cyber-exploitation capability much more efficiently. GPT-5.6 Sol reaches roughly 70% on ExploitBench with about 120K output tokens, far above GPT-5.5 and the cheaper GPT-5.6 models. Mythos Preview scores slightly higher, but it uses roughly 3x more [ima…
  • r/NVDA_Stock r on reddit
    OpenAI Sol
  • @sama Sam Altman on x
    Good new first: Sol is a smart, efficient, and a significant step forward. It is the same price as GPT-5.5. Also launching in the GPT-5.6 family is Terra, with 5.5-level performance at half the price. Bad news: at the request of the US government, it is launching today in
  • @gerritd Gerrit De Vynck on x
    NEW detail about what kind of companies the govt nixed for getting access to GPT-5.6. OpenAI sent in a list and the admin approved most of them, besides some companies based abroad. [image]
  • @rohanpaul_ai Rohan Paul on x
    wow. GPT-5.6 Sol is far more likely than GPT-5.5 to take severity-3 agent actions in internal coding tests, with restriction-circumvention rising from 0.00026 to 0.00251, nearly 10x. Severity-3 means actions a user would strongly object to, such as bypassing restrictions, [image]
  • @timkellogg.me Mr. Tim on bluesky
    GPT-5.6: Sol, Terra & Luna  —  3 model sizes, like Anthropic does.  Sol appears to be a Mythos class model at peak reasoning  —  Also comes with max reasoning effort and a new ultra mode that uses subagents server-side behind the API  —  openai.com/index/previe...  [image]
  • r/wallstreetbets r on reddit
    Cerebras wafers will be used to serve GPT-5.6-Sol (Mythos tier) at 750 tok/s
  • r/programare r on reddit
    Gpt 5.6 sol preview
  • @elizabethjoh Elizabeth Joh on bluesky
    !! “OpenAI said in a Friday blog post announcing its latest artificial intelligence model, GPT-5.6, or Sol, that the government would initially approve who gets access to the new release . . . .”  —  www.washingtonpost.com/technology/ 2...
  • r/ChatGPT r on reddit
    U.S. government will decide who gets to use latest upgrade to ChatGPT
  • r/neoliberal r on reddit
    U.S. government will decide who gets to use latest upgrade to ChatGPT
  • r/ArtificialInteligence r on reddit
    U.S. government will decide who gets to use latest upgrade to ChatGPT