Claude Opus 4.8 and Sonnet 5 seem worse at tool calling than older models, likely as post-training optimized them primarily for Claude Code-like environments
A very strange Pi issue sent me down a rabbit hole over the last two days. The short version is that newer Claude models sometimes call Pi's edit tool …
Context & Ripple Effects
Claude’s recent releases have been positioned around coding, computer use, instruction following and agentic work, building on earlier support for alternating reasoning with tool use. Sonnet 5 was specifically presented as approaching Opus 4.8 performance at lower prices.
This report challenges the portability of those gains: behavior optimized for a Claude Code-like harness may not translate cleanly to Pi’s editing tool or other tool interfaces. It also lands amid user complaints about perceived performance changes in earlier Claude and Claude Code versions.
First-order effects
- Pi users and developers may see less reliable edit-tool behavior from Claude Opus 4.8 and Sonnet 5 than from older models, requiring model-specific fallbacks or prompt and tool-interface adjustments.
- Anthropic faces a concrete compatibility signal: models marketed for agentic work can perform unevenly when tool semantics differ from the environment emphasized in post-training.
Second-order effects
- Tooling vendors will have stronger reason to benchmark models against their own schemas and execution loops rather than treating a model’s coding or agentic score as portable across harnesses.
- Teams deploying Claude across multiple agents may keep older models in production for particular tool workflows, reducing the simplicity of upgrading to the newest, lower-cost release.
Third-order effects
- If similar reports recur, agentic-model competition will shift from broad claims of tool-use capability toward compatibility with specific harnesses, tool contracts and evaluation suites.
- Post-training may increasingly create ecosystem-specific strengths: model providers and agent-platform builders could gain leverage by controlling the environments used to train and measure tool use.
The trend: This is one data point in the shift from general-purpose model releases toward agent systems whose real-world performance depends on the fit between model post-training and the surrounding tool harness.