Anthropic details three infrastructure bugs that intermittently degraded Claude's responses between August and early September, and explains how it fixed them
This is a technical report on three bugs that intermittently degraded responses from Claude. Below we explain what happened …
Anthropic
Context & Ripple Effects
This disclosure sits in a recurring record of Claude reliability and quality problems, including a later incident in which Anthropic said it had fixed three Claude Code quality causes and a separate period of elevated errors across Claude surfaces followed by a stated fix. The repeated pattern makes root-cause transparency consequential for users deciding whether degraded output is a model limitation, a product-policy change, or an operational fault.
First-order effects
Anthropic’s fixes should remove the identified intermittent failure modes for affected Claude users and give customers a more concrete explanation for the degraded responses.
The report distinguishes infrastructure-caused quality degradation from intentional model behavior, directly addressing a source of uncertainty for users evaluating Claude output.
Second-order effects
Enterprise and developer users may place greater weight on monitoring, validation, and fallback workflows when Claude quality shifts, rather than treating output changes as solely model-level behavior.
The disclosure creates a benchmark for competing AI providers: operational quality claims increasingly need incident explanations and remediation details, not just aggregate availability signals.
Third-order effects
As AI systems become embedded in production work, reliability management will extend from uptime to output quality: providers may be judged on whether they can detect, diagnose, and communicate silent degradations.
If such episodes persist across the sector, deployment governance will likely treat model quality, serving infrastructure, and product configuration as a single operational risk surface.
The trend: Frontier AI providers are moving toward operational accountability for output quality, not merely service availability.
@_sholtodouglas it's nice to see this being discussed sholto, but doesn't make it any less disappointing. many of us paying $200 a month for something that didn't work as well as it did 1 month ago. the fact despite the loud chorus of complaints anthropic said nothing. maybe it w…
We've published a detailed postmortem on three infrastructure bugs that affected Claude between August and early September. In the post, we explain what happened, why it took time to fix, and what we're changing:
@_sholtodouglas too late :( codex's timing and cursor price changes really made it very easy to break from using claude code and claude on cursor after paying both anthropic and cursor for more than a year. probably not gonna touch claude unless next version is insanely better
one of the best postmortem activity of the year. this shows where the SRE job is heading and why will remain relevant. and I have to say that's pretty damn great where we are traveling to. kudos to @AnthropicAI (I'm sure @tmu was here ;)
Anthropic full postmortem — long context routing bug — output corruption bug — approx top-k bug on TPU — plus: why it was so difficult and slow to fix — this is going to be a thread — www.anthropic.com/engineering/ ...