Claude users accuse Anthropic of degrading Claude Opus 4.6's and Claude Code's performance; Anthropic staff publicly deny it degrades models to manage capacity
From Its Own FansMaria Garcia /Implicator.ai:Anthropic Ships Claude Code Routines, Cloud Automations That Run Without Your MacLeila Sheridan /Inc.com:Users Say Anthropic's Claude Is Getting Worse. A Quiet Change May Be to BlameCraig Hale /TechRadar:‘Claude cannot be trusted to perform complex engineering tasks’: AMD AI head slams Anthropic's coding tool after months of frustration
VentureBeatCarl Franzen
Context & Ripple Effects
Anthropic had recently made Claude Code generally available and then tightened session limits during peak periods as Claude’s popularity put pressure on available compute. That sequence makes performance consistency—not just model capability—a central issue for users relying on the product.
The dispute is especially consequential because it pairs user reports with a public denial from Anthropic staff, while a prominent AMD AI executive has separately criticized Claude Code’s reliability on complex engineering work.
First-order effects
Claude Opus 4.6 and Claude Code users face immediate uncertainty over whether observed quality changes stem from product behavior, usage limits, or other service conditions; Anthropic must defend confidence in its service without conceding intentional degradation.
For engineering teams using Claude Code, public reliability criticism raises the practical cost of trusting the tool for complex tasks and increases pressure to validate outputs and maintain fallback workflows.
Second-order effects
Capacity-management choices such as peak-hour limits become harder to separate, in customers’ minds, from perceived quality reductions; providers will be pushed to communicate service constraints and product changes more clearly.
Competing coding-assistant and model providers gain an opening to compete on predictable performance and transparent limits, rather than only headline model capability.
Third-order effects
As AI coding tools move from experiments into recurring engineering workflows, reliability under constrained inference capacity may become a differentiator as important as benchmark performance.
If complaints of this kind recur, enterprise buyers are likely to evaluate AI services on task-level consistency, controls, and operational transparency—not merely access to a named model.
The trend: This is part of a broader shift in which inference capacity and service-policy decisions increasingly determine the real-world usefulness of AI models and agents.
This is false. We defaulted to medium as a result of user feedback about Claude using too many tokens. When we made the change, we (1) included it in the changelog and (2) showed a dialog when you opened Claude Code so you could choose to opt out. Literally nothing sneaky abou…
AMD Senior AI Director confirms Claude has been nerfed. She analyzed Claude's session logs from Janurary to March: > median thinking dropped from ~2,200 to ~600 chars > API requests went up 80x from Feb to Mar. less thinking and failed attempts meaning more retries, burning more …
SOMEONE ACTUALLY MEASURED HOW MUCH DUMBER CLAUDE GOT. THE ANSWER IS 67%. the data shows Opus 4.6 is thinking 67% less than it used to. anthropic said nothing until the numbers went public. then suddenly Boris Cherny (creator of Claude Code) shows up on the GitHub issue. users ar…
Despicable clout chasing. They tested Opus today on 30 tasks, previous Opus 4.6 score was on just *6* tasks. DIFFERENT BENCHMARK 6 tasks in common results: 85.4% score today vs. 87.6% prev. Swing is mostly from a *single* fabrication without repeats - easily statistical noise [im…
this is not a new argument btw. Lex Fridman asked Dario Amodei the same question, “is Claude getting dumber?” back in 2024 for Sonnet 3.5. It's just surprising there's still not an open, transparent daily/weekly performance report of Codex & Claude similar to uptime services [vid…
I think the anthropic people are gaslighting us, sidestepping questions and answering a different question in lawyer speak brand trust erosion speedrun any%
basically: anthropic sneakily turned down how hard claude thinks before editing code, changed the default from “high” to “medium” effort, and hid the reasoning from session logs. all without telling users. an amd director had 7k sessions of telemetry to prove the degradation
I completely agree with her assessment that “Claude has regressed to the point it cannot be trusted to perform complex engineering”. There are two problems with building businesses on top of frontier models. The first is the cost and value capture. The frontier labs are aiming
is the intended meaning that changing the settings in a way that degrades the output isn't “degrading the model” (the model is the same, the settings are just tokens in the context!) or that it is “degrading the model” but that it was done for some other reason?
It reminded me of the early days of the Internet, when bandwidth was shared with neighbors. Then the upstairs neighbor would try to find the kids taking up all the bandwidth downloading porn or pirated files.
@edzitron Sometimes they make tweaks to Claude Code's harness to try to improve something, and there's a weird byproduct of it reducing overall quality of output and burning tokens It happens, but mostly with Claude over the other models. Claude seems more sensitive to stuff like…
Anthropic allegedly doesn't degrade its models intentionally, and yet they seem to be the only lab that experiences this degradation issue going back to Claude 3 out of every other SOTA lab. Are the models inherently unstable and black box, or is there something more to this?
I have seen scattered reports of Claude burning more tokens, and it does seem like token burn increased on openrouter in this period too, wonder if 4.6 is also part of it?
They're killing Claude on purpose-200,000 tasks prove it. Someone analyzed the data. Claude isn't “having a bad day”. It's being systematically throttled. The Breakdown: > Stopped verifying: Used to check 6 times before acting. Now it glances once and wings it. > Sandbagging: [im…
Claude being nerf'd and agents being exiled from the $200/mo plan are very consistent behavior if Anthropic is going public soon and will have finances/margin intensely scrutinized...
So the “Anthropic secretly nerfed Claude” narrative looks like fake news. According to an Anthropic dev, they mainly stopped showing thinking summaries by default for latency, which distorted the measurements. That is not the same thing as secretly downgrading the model. [image]
This is not enough. Claude and Anthropic need to explain all the measured increases in hallucination and degration as observed by AMD. You can't just handwave this away.
@Hesamation boris responded to this in depth in the issue- it's mostly just that we stopped showing thinking summaries for latency (you can opt-in to showing it) which was affecting the thinking measurement in the post https://github.com/...
you can run the nerfing play once, maybe twice. but anthropic will silently degrade production models to farm failure data every time. and this is where it stops being a clever strategy and starts being a trust problem.
Not enough compute is the correct take IMO. Which has quite a lot more implications if you think about it and play that out to its logical conclusion. Software still burdened by hardware's inability to keep up. Maybe software is 2-3 years ahead of hardware?
Imagine being one of those CEOs who laid off thousands over “AI efficiency” only for the AI to get dumber than a pile of bricks weeks later. — Given how tokens work, you're paying more for a worse product. It now takes 5 minutes to be wrong which took 30 seconds for a right an…