Data from 300K+ pull requests shows OpenAI is catching up to Anthropic in AI coding: Codex has a 74.3% success rate vs. Claude Code's 73.7% in code approvals
OpenAI's effort to catch up to Anthropic in code-generating artificial intelligence seems to be working.
Context & Ripple Effects
This narrows a coding-model rivalry that had already pushed OpenAI to improve ChatGPT’s coding capabilities in response to Claude, according to prior coverage of the OpenAI–Anthropic competitive response.
The comparison matters because coding agents are moving from a feature benchmark toward an internal work tool: OpenAI later reported broad employee adoption of Codex across its workforce, while Anthropic reported Claude-authored code making up most of merged code in its own codebase.
First-order effects
- The pull-request analysis gives Codex a narrow lead on this approval-based measure, challenging Claude Code’s apparent advantage in a highly visible coding-agent use case.
- Engineering teams evaluating the two products gain a directly comparable signal, but the small gap makes local testing and workflow fit important before treating either as decisively ahead.
Second-order effects
- A near-tie increases pressure on both vendors to differentiate beyond raw code acceptance—through developer workflow, reliability, deployment and distribution rather than a single benchmark result.
- Buyers are likely to scrutinize the cost and review burden of accepted code more closely, since comparable approval rates alone do not establish which tool delivers more useful work per engineering dollar.
Third-order effects
- If coding agents continue to produce a growing share of merged code, vendor competition will increasingly center on the systems that govern review, integration and accountability—not just model output quality; Anthropic’s reported high share of Claude-authored merged code illustrates that shift.
- Approval-based production data could become a more consequential evaluation standard than demo-style coding tests, though its usefulness will depend on whether results hold across repositories, teams and review policies.
The trend: AI coding is shifting from model-level capability races toward competition over measurable, production-grade software-delivery outcomes.