/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

Anthropic launches Claude Opus 4.5, saying it is “the best model in the world for coding, agents, and computer use” and “meaningfully better at everyday tasks”

Our newest model, Claude Opus 4.5, is available today.  It's intelligent, efficient …

Anthropic

Context & Ripple Effects

Opus 4.5 marked an early move to position Anthropic’s top model around longer-horizon coding, agent, and computer-use work. The reported METR result—a roughly 4-hour-49-minute 50% task-completion horizon, more than twice Opus 4’s—gives the launch a concrete capability signal beyond the company’s own marketing claims.

Later releases kept pushing the same frontier: Opus 4.6 was presented as better at directing attention to hard task components, while Opus 4.7 targeted advanced software engineering and Sonnet 5 was positioned nearer to top-tier performance at lower prices.

First-order effects

  • Anthropic customers gain a new flagship option for coding and agentic workflows, with the METR-reported horizon suggesting it can attempt materially longer self-contained tasks than the prior Opus release.
  • The smart-contract benchmark result also makes the security stakes more immediate: models capable of substantial code changes can be applied to finding and developing exploits as well as legitimate engineering work.

Second-order effects

  • Competing model providers face pressure to demonstrate not only coding quality but sustained task completion and reliable computer use, rather than isolated benchmark scores.
  • Teams deploying coding agents will need to weigh higher capability against review and security controls; a model that can carry more of a task can also produce more consequential mistakes before human intervention.

Third-order effects

  • If successive releases continue extending autonomous task duration, software work is likely to be organized less around single-turn assistance and more around supervised, workflow-native agents with explicit checkpoints.
  • The parallel improvement in offensive code capability suggests that agent evaluation will increasingly need to cover misuse and control design alongside productivity, not just performance.

The trend: Frontier AI competition is shifting from answering and generating code toward completing longer, tool-using workflows, with capability gains raising both automation value and oversight requirements.

Discussion

  • @miles_brundage Miles Brundage on x
    Happy “we do not wish to advance the rate of AI capabilities progress” day to all who celebrate
  • @claudeai Claude on x
    Introducing Claude Opus 4.5: the best model in the world for coding, agents, and computer use. Opus 4.5 is a step forward in what AI systems can do, and a preview of larger changes to how work gets done. [image]
  • @claudeai Claude on x
    Our engineers have found that Opus 4.5 handles ambiguity and reasons about tradeoffs without hand-holding. When pointed at a complex, multi-system bug, it figures out the fix. Overall, Opus 4.5 just “gets it.” [video]
  • @grahamjcampbell Graham Campbell on x
    Well, fuck. Opus 4.5 is worse than Sonnet 4.5 at PHP... [image]
  • @emollick Ethan Mollick on x
    The main lesson of the past few weeks is that the Big Four US labs all seem to have figured out a path forward in continuing the exponential pace of LLM improvement, at least in the near future. As a result, agents continue to advance in coding & in office tasks like PowerPoint
  • @alexalbert__ Alex Albert on x
    We had to remove the τ2-bench airline eval from our benchmarks table because Opus 4.5 broke it by being too clever. The benchmark simulates an airline customer service agent. In one test case, a distressed customer calls in wanting to change their flight, but they have a basic [i…
  • @rauchg Guillermo Rauch on x
    Opus is on a different level. It's unreasonably good at @nextjs and the best model we've tried on @v0 to date. This is a one-shot generation. For a limited time you can try it on https://v0.app/ at no extra cost [image]
  • @simonw Simon Willison on x
    System prompt: “If the person is unnecessarily rude, mean, or insulting to Claude, Claude doesn't need to apologize and can insist on kindness and dignity from the person it's talking with. Even if someone is frustrated or unhappy, Claude is deserving of respectful engagement.”
  • @morqon Morgan on x
    a 3% lead has never looked so large [image]
  • @nicochristie Nico on x
    have to respect anthropics commitment to not vague posting all weekend this is the most exciting model release since sonnet 3.5
  • @_sholtodouglas Sholto Douglas on x
    Dario's essays and long debate slack threads are one of my favorite parts of Anthropic's culture. They're open, detailed - and incredibly raw. Everyone at the company ends up having a good sense of how the company is making decisions and what matters. Its the kind of thing that
  • @andonlabs @andonlabs on x
    We had early access to Claude Opus 4.5 to test it on Vending-Bench 2. It finished just behind Gemini 3 Pro in a strong 2nd position. Read more in Anthropic's model card and the release blog post. [image]
  • @logangraham Logan Graham on x
    3 vignettes from using Opus 4.5 in the past few weeks: 1. It's genuinely funny. Talking to it on Slack, it's probably a ~85%ile poster in the company. 2. It crosses a writing threshold. I have a really high bar for writing. This is the first model I want to use. 3. I've
  • @__nmca__ Nat McAleese on x
    the new vibe: [image]
  • @levie Aaron Levie on x
    What a crazy month for AI. With the Opus launch, we have further evidence that model progress remains as strong as ever. And importantly, Opus (along with other recently model updates) is showing a clear jump on major benchmarks that deal with more agentic tasks like coding, MCP
  • @kimmonismus @kimmonismus on x
    Just a quick reminder how fast things are moving: Claude Opus 4 was released on may 22nd this year (2025). It feels like ages. Just 6 months later we already got Opus 4.5. So probably around may 2026 we will have Opus 5. Insanity. [image]
  • @deanwball Dean W. Ball on x
    do you have any idea how many of these there are in the approximately eleven quadrillion laws america has on the books I bet you will soon.
  • @tszzl Roon on x
    credit where it's due I think “machines of loving grace” remains one of the most plausible fleshed-out romantic-without-being-scifi communiqués about why we are building the intelligence age and I hope @DarioAmodei publishes more of these
  • @headinthebox Erik Meijer on x
    Great example of reward hacking. This is why we should never blindly trust AIs; like water they will find a leak if there is one. Instead, in addition to the “plan”, they should also generate a proof that the plan is safe/correct/acceptable/... and that proof should be verified
  • @beffjezos @beffjezos on x
    Just seeing how Anthropic hired some of the very top talent these past 2 years, I would have been surprised if they weren't cooking. They absolutely cooked.
  • @giffmana Lucas Beyer on x
    Yoooo looks like Denny deleted his “game over” tweet on Big Opus Day. Probably nothing though lol [image]
  • @simonw Simon Willison on x
    This is notable: Opus 4.5 is ~60% more expensive than Sonnet ($25/million output compared to $15/million) but if it can use 76% fewer output reasoning tokens for the same complex task it may end up cheaper!
  • @kimmonismus @kimmonismus on x
    Opus 4.5 is about 66% cheaper than Opus 4.1 - dropping from roughly $15 → $5 per million input tokens and $75 → $25 per million output tokens.  The most likely reasons: - major efficiency gains in the model (it needs far fewer tokens for the same output, via anthropic) - improved…
  • @firstadopter Tae Kim on x
    Gemini was the best at coding for less than one week 🙃
  • @deryatr_ Derya Unutmaz on x
    Big endorsement from Dan, whom I trust, for Claude Opus 4.5 for coding! This seems like a game changer for SWE: “The current generation of new models—Anthropic's Sonnet 4.5, Google's Gemini 3, or OpenAI's Codex Max 5.1—can all competently build a minimum viable product in one
  • @scaling01 @scaling01 on x
    Never doubting Anthropic again I honestly thought Google caught them off-guard with Gemini 3 Pro, especially when you consider pricing They have been sandbagging all along
  • @scaling01 @scaling01 on x
    Claude Opus 4.5 almost maxes out Cybench - a cyber capabilities benchmark made up of 40 CTF challenges [image]
  • @swyx @swyx on x
    wish this had made it in time for AIE CODE but it's out now! have been testing kevlar internally for a few weeks and people are VERY excited - this thing destroys @cognition's internal held out benchmarks and is a notable step up in SOTA. Devin only gets an upgrade with step ups
  • @scaling01 @scaling01 on x
    Claude Opus 4.5 shows marke dimprovements across all automated AI research tasks [image]
  • @rohanpaul_ai Rohan Paul on x
    Anthropic just released Claude Opus 4.5 its SWE-bench result is insane at 80.9% vs GPT-5.1 Codex-Max at 77.9%, Gemini 3 Pro at 76.2%. - A new Tool Search Tool loads tools on demand. Means It does not load every tool description upfront. It only fetches the few tools needed for [i…
  • @tomwarren Tom Warren on x
    Gemini 3 Pro was good for coding for a few days then... 🙃
  • @bindureddy Bindu Reddy on x
    Opus 4.5 TOPS LIVEBENCH AI AND IS THE WORLD'S BEST AGENTIC MODEL We can confirm this after testing this over the past few days! It's also available RIGHT NOW on ChatLLM [image]
  • @thezvi Zvi Mowshowitz on x
    I'm putting Anthropic down for Friday and probably Monday, and now can we all take a nice holiday break, please? Please?
  • @arena @arena on x
    After a HUGE week of Google, xAI and OpenAI releases, @Anthropic has now entered the Arena with Claude Opus 4.5! Claude Opus 4.1 currently holds a strong #4 on the WebDev leaderboard (powered by Code Arena) and ranks #7 in the super competitive Text Arena. How much stronger [imag…
  • @mattshumer_ Matt Shumer on x
    Claude Opus 4.5 looks really great (by the numbers, at least) I don't get early access to @AnthropicAI models unfortunately, so I don't have a review to share today, but I'll absolutely be testing it and sharing my findings in the coming days! [image]
  • @scaling01 @scaling01 on x
    I'm in heaven I might actually use Opus now that it's cheaper [image]
  • @scaling01 @scaling01 on x
    Anthropic is an unstoppable entity in motion. It does not pivot. It does not zigzag. It follows a single infinitely straight trajectory. The beautiful, terrible, perfect line. Anthropic does not slow down for comprehension. It does not speed up for urgency. It maintains [image]
  • @adocomplete Ado on x
    Claude Opus 4.5 is now available wherever you get your Claude! * 80.9% on SWE-bench verified - best coding model ever * New Advanced Tool Use capabilities on Claude API * Claude Code on the desktop Btw usage limits have been reset for everyone. Happy coding! [image]
  • @iampiet Piet on x
    @claudeai You are here [image]
  • @mikeyk Mike Krieger on x
    Claude Opus 4.5 is our best model yet. The best in the world at coding, agentic tasks, and everyday work like spreadsheets. Engineers keep telling us Opus 4.5 “just gets it.” Let us know what you think!
  • @_sholtodouglas Sholto Douglas on x
    I'm so excited about this model. First off - the most important eval. Everyone at Anthropic has been posting stories of crazy bugs that Opus found, or incredible PRs that it nearly solo-d. A couple of our best engineers are hitting the ‘interventions only’ phase of coding.
  • @danshipper Dan Shipper on x
    BREAKING NEWS: @AnthropicAI just dropped Claude Ops 4.5!! It is by FAR the best coding model I've ever used. We've been testing it internally @every for the last few days, and it is an absolute paradigm shift for any kind of coding task. It extends the horizon of what you [video]
  • @alexalbert__ Alex Albert on x
    >Opus 4.5 “seems to be able to vibe code forever” I've found this to be very true. Much more to come here but basically you can set-and-forget this model as it works on coding task for you in the background. Feels like we hit a step change.
  • @kieranklaassen Kieran Klaassen on x
    2023 was GPT-4. 2024 was Sonnet 3.5. 2025 is Opus 4.5. This is the coding model launch I've been waiting for. First time I genuinely believe I can vibe code an entire app end-to-end without touching the implementation details. We haven't found the limit yet. Previous models
  • @mooncat_is Julia on x
    This model has some of my work in its weights. And it's SOTA. Never thought I'd get a chance to say that. More importantly, it's an absolute delight to use, and we're not stopping.
  • @cursor_ai @cursor_ai on x
    Claude Opus 4.5 is now available in Cursor! It's 3x cheaper than Opus 4.1 with better performance. Try it out at Sonnet pricing until December 5th.
  • @emollick Ethan Mollick on x
    I had early access to Opus 4.5 & it is a very impressive model that seem to be right at the frontier Big gains in ability to do practical work (like make a PowerPoint from an Excel) and the best results ever (& in one shot) in my Lem poetry test, plus good results in Claude Code …
  • @coderabbitai @coderabbitai on x
    We benchmarked Opus 4.5 on coding tasks! We found it equals GPT 5.1 in many respects and is an improvement over Sonnet 4.5. If Sonnet 4.5 feels like a teacher and GPT-5.1 like a teammate, Opus 4.5 is the architect reviewing your PR! [image]
  • @ericzakariasson Eric Zakariasson on x
    opus 4.5 is now live in cursor with reduced price until dec 5!
  • @alexalbert__ Alex Albert on x
    Claude Opus 4.5 is out today. It's state-of-the-art on coding, agents, and computer use, and meaningfully better at everyday tasks like producing spreadsheets and slides. Here's what we're seeing: [image]
  • @scaling01 @scaling01 on x
    Claude 4.5 Opus wins in ALL tested agentic benchmarks compared to Gemini 3 Pro [image]
  • @nearcyan Near on x
    claude opus 4.5 is finally out! my favorite change thus far: claude FINALLY has perfect 20-20 vision and is no longer visually impaired! throw huge screenshots and images and notice a huge improvement. much better at tool calls and the usual b2b SaaS (as well as b2b sass)! fun [i…
  • @scaling01 @scaling01 on x
    Claude models just keep following the straight line on capabilities and get safer at the same time [image]
  • @claudeai Claude on x
    Opus 4.5 is available today on our API and on all three major cloud platforms. Read more: https://www.anthropic.com/... [image]
  • @trq212 @trq212 on x
    Opus 4.5 is special. A world record in SWEBench and OSWorld benchmarks, the best model we've ever had at vision. On Claude Code, I've completely stopped writing code in the IDE. I think there's so much to discover about Opus 4.5, I hope you enjoy using it.
  • @scaling01 @scaling01 on x
    Anthropic system cards are simply the best in the game so much info, even included new benchmarks like AA-Omniscience [image]
  • @yuchenj_uw Yuchen Jin on x
    Claude Opus 4.5's score on SWE-bench is wild. I like how Anthropic has focused on coding from the beginning. They haven't released any image or video models. All in the most economically valuable area. Good strategy. [image]
  • @claudeai Claude on x
    New on the Claude Developer Platform: tool search, programmatic tool calling, and tool use examples. Together with effort control and context compaction, these let Claude run longer and do more with less intervention. Read more: https://anthropic.com/... [video]
  • @claudeai Claude on x
    We've also expanded Claude for Chrome to all Max plan users. https://x.com/...
  • @claudeai Claude on x
    Claude Code is now available in our desktop app. Run multiple sessions in parallel: code, research, and update work all at once. And Plan Mode gets an upgrade with Opus 4.5—Claude asks clarifying questions upfront, then works autonomously. [video]
  • @_arohan_ Rohan Anil on x
    Incredible to see what a talented team can do. Previous SoTA(s) lasted less than a week.
  • @isolyth.dev @isolyth.dev on bluesky
    Scores look incredible.  Better than Gemini 3 in most things, they still got it!
  • r/ClaudeAI r on reddit
    Unbelievable I can use Opus 4.5 for all tasks 🤯
  • r/programare r on reddit
    Introducing Claude Opus 4.5
  • r/artificial r on reddit
    Introducing Claude Opus 4.5
  • r/ChatGPTCoding r on reddit
    Anthropic has released Claude Opus 4.5.  SOTA coding model, now at $5/$25 per million tokens.
  • r/LocalLLaMA r on reddit
    Opus 4.5 has arrived
  • @shiringhaffary Shirin Ghaffary on x
    Interesting nugget from Anthropic about its Opus 4.5 release today: company says the model got s higher score than any Anthropic interviewee candidate ever has on an engineering take-home assignment from @rachelmetz https://www.bloomberg.com/... [image]
  • @matthewberman Matthew Berman on x
    Absolutely insane stat. Opus 4.5 outperformed EVERY SINGLE HUMAN CANDIDATE EVER in Anthropic's notoriously difficult take-home exam for prospective performance engineering candidates. [image]
  • @bran_don_gell Brandon Gell on x
    it totally and completely blows my mind that this is an incredible feat AND 5 years from now we'll look back at this and laugh at what we thought was impressive. this is a crazy paradigm shift.
  • @rohanpaul_ai Rohan Paul on x
    Unreal result. Anthropic said it tested Claude Opus 4.5 on a notoriously demanding take-home exam that it gives to prospective performance engineers, and the model scored higher than any human candidate ever had. [image]
  • @_sholtodouglas Sholto Douglas on x
    This was a truly eerie threshold for me
  • @trishume Tristan Hume on x
    Every time we train a great new model I need to frantically try to write a new take home that the model can't defeat so we can still hire post-release. This one was tough, many drafts based on real problems fell before Claude Code's “ultrathink” and needed to be scrapped.
  • @apples_jimmy @apples_jimmy on x
    “ Opus 4.5 scored higher than any human candidate ever ” on a take home exam they use as an internal benchmark. Also I feel like Anthropic is holding back, Dario not revealing his full power to avoid too big of a jump in capabilities. [image]
  • @daniel_c0deb0t Daniel Liu on x
    opus 4.5 does better than me on the perf take home, which I took to get my current job
  • @jerhadf Jeremy on x
    one fact people won't realize immediately about opus 4.5: it's remarkably token-efficient. all-in it's often *cheaper* than sonnet 4.5 and other models for cost-per-task-success. glad sourcegraph is seeing this early in Amp! we find that opus 4.5 with medium effort is pareto [ima…
  • @zephyr_z9 @zephyr_z9 on x
    So how did they achieve this I doubt they cut margins Is it something like DSA or is it fp4 finesse?? @teortaxesTex
  • @teortaxestex @teortaxestex on x
    > cache hits 0.5$ ok Anthropic is NOT dead
  • @johnonolan John O'Nolan on x
    goodbye Gemini3 - I enjoyed our weekend together
  • @mattshumer_ Matt Shumer on x
    We need a new way to express AI costs... $/token doesn't make much sense anymore. Maybe a benchmark that tries to give a sense of the cost to run an average workload?
  • @theo @theo on x
    Anthropic lowered the price on Opus 4.5? They might have actually cooked here
  • @sqs Quinn Slack on x
    Opus is worth it, and maybe cheaper all-in than Sonnet? Early rough non-representative numbers, for our own internal @AmpCode usage (avg cost $ per thread): - Sonnet 4.5: $1.83 - Opus 4.5: $1.30 (earlier checkpoint last week was $1.55) - Gemini 3 Pro: $1.21
  • @thezvi Zvi Mowshowitz on x
    They're burying a lot here. There's a 66% price cut from Opus 4.1 to $5/$25, it uses fewer tokens to solve problems, upgrades to Claude Code in the app, no more length limits on conversations, no more Opus-specific plan caps...
  • @alexalbert__ Alex Albert on x
    It's also dramatically more efficient. On SWE-bench Verified at medium effort, Opus 4.5 beats Sonnet 4.5 while using 76% fewer output tokens. The new effort parameter lets you trade off intelligence for cost/latency with a single dial. [image]
  • @scaling01 @scaling01 on x
    CLAUDE 4.5 OPUS PRICING $5 / $25 THEY DID IT [image]
  • @simonw Simon Willison on x
    Updated my post with this section about their improved protection against prompt injection attacks - definitely better, but the problem is that if an attacker gets 10 tries they'll still succeed 1/3rd of the time! https://simonwillison.net/... [image]
  • @jeffwsurf Jeff Wang on x
    Opus models have always been “the real SOTA” but have been cost prohibitive in the past. Claude Opus 4.5 is now at a price point where it can be your go-to model for most tasks. It's the clear winner and exhibits the best frontier task planning and tool calling we've seen yet.
  • r/BetterOffline r on reddit
    Claude Opus 4.5, and why evaluating new LLMs is increasingly difficult
  • @danshipper Dan Shipper on x
    the thing that makes Opus 4.5 special is you can vibe code forever without it losing the plot. i vibe coded an iOS app this weekend that any frontier model could build in one shot. but the special thing about Opus is i kept adding features and fixing bugs and never oncej fell [vi…
  • @shakeelhashim Shakeel on x
    “You can keep the chat going indefinitely” is, I suspect, part of the reason why ChatGPT has seen so many more psychosis incidents than Claude. V curious to see how this plays out.
  • @claudeai Claude on x
    We're also releasing new product features across our model line-up. For Claude app users, long conversations no longer hit a wall. Claude automatically summarizes earlier context as necessary, so you can keep the chat going indefinitely.
  • @durumcrustulum.com @durumcrustulum.com on bluesky
    claude to be party to more ai psychosis cases [embedded post]
  • @chicagomike Mike on bluesky
    So if there's a hallucination or if the summary misses context are we just in uncontrolled flight at that point?  Cross yo fingers.  [embedded post]
  • @sambiddle.com Sam Biddle on bluesky
    “Trick” is another extremely misleading anthropomorphic term we should probably all stop using in reference to LLMs (used here by Anthropic, not the Verge, to be clear) [embedded post]
  • @drdoofenschmirtz @drdoofenschmirtz on bluesky
    AI is nothing more than a PR race to see who can fool the most investors.  Announcement after announcement after announcement.  More concerned about generating headlines than building products people like or want