Some developers say GPT-5 excels at technical reasoning and planning coding tasks and is cost-effective, but Claude Opus and Sonnet still produce better code
Software engineers are finding that OpenAI's new GPT-5 model is helping them think through coding problems—but isn't much better at actual coding.
Context & Ripple Effects
The report complicates the recent arc around GPT-5: sources had described improved practical software-engineering performance, while OpenAI positioned the release as a routed system combining efficient and reasoning models. Developer feedback suggests those design gains may be more visible in problem framing than in final code output.
That distinction matters because coding teams buy a workflow, not a benchmark category: planning, implementation, review, and debugging can be assigned to different models when their relative strengths diverge.
First-order effects
- Developers can use GPT-5 for technical reasoning and task planning while retaining Claude Opus or Sonnet for code generation where they judge its output stronger.
- OpenAI’s coding proposition is evaluated on cost-effective assistance and planning quality, rather than on a clear claim of superior working-code generation.
Second-order effects
- Teams comparing model vendors are likely to measure cost per completed engineering task, not merely coding benchmark or token-price results; this favors task-level routing over a single-model standard.
- Claude’s reported code-output advantage gives it a concrete retention point in developer workflows, even where GPT-5 is attractive for earlier stages of the same task.
Third-order effects
- If this division persists, AI coding stacks will become more modular: reasoning, planning, generation, and verification may be sourced separately and switched as model performance changes.
- Competition will increasingly center on workflow fit and reliable end-to-end outcomes, making model routing and workload portability more consequential than broad claims of coding capability.
The trend: AI coding is shifting from a race for one best model toward workflow-level orchestration based on the cost and quality of each stage of engineering work.