A look at GPT-4.5's claimed performance, including on coding benchmarks, where it matches or outperforms GPT-4o but falls short of OpenAI's Deep Research
Bigger, smarter, and more powerful Samantha Kelly / CNET : OpenAI Says Its New ChatGPT 4.5 Has Better Emotional Intelligence Cristina Criddle / Financial Times : OpenAI reveals GPT-4.5 amid flurry of new AI model releases Markus Kasanmascheff / WinBuzzer : OpenAI Announces GPT-4.5, Its Largest LLM To Date Bluesky: @fintechdad : “On several AI benchmarks, GPT-4.5 falls short of newer AI “reasoning” models from Chinese AI company DeepSeek, Anthropic, and OpenAI itself” hmmmm. [embedded post] Mastodon: Bret Carmichael / @bretcarmichael@mastodon.social : AI competitors continue to merge, and competition is making them race to release new models more quickly. The individual versions are trending to meaningless, as users just use the best things available through their default provider. — https://techcrunch.com/... #AI #OpenAI #ChatGPT X: Aaron Levie / @levie : Box tested GPT-4.5 vs. GPT-4o in Box AI with 20,000+ data fields being pulled from enterprise content (like important details in a contract). We found a 19 pt improvement in single shot extraction. This is a huge improvement for any mission critical enterprise workflow. [image] Jeremy Howard / @jeremyphoward : This graph (from @enricoros, based on @paulgauthier work) show's GPT-4.5's perf on (IMO) the most practical & important eval: Aider Polyglot. If this is accurate, OpenAI is in *deep* trouble! Their crown jewel, the latest GPT, is *500x* more $$$ than DeepSeek v3, yet worse! [image] Eli Lifland / @eli_lifland : And now I lengthen my timelines, at least if my preliminary assessment of GPT-4.5 holds up. Not that much better than 4o (especially at coding, and worse than Sonnet at coding) while being 15x more expensive than 4o, and 10-25x more expensive than Sonnet 3.7. Weird. [image] Jeremy Howard / @jeremyphoward : This is specifically for coding, which is by far the most common use of LLMs, and is highly correlated with how good they are at other important tasks. Gary Marcus / @garymarcus : Hot take: GPT 4.5 is mostly a nothing burger. GPT 5 is still a fantasy. • Scaling data and compute is not a physical law, and pretty much everything I have told you was true. • All the bullshit about GPT-5 we listened to for the last couple years: not so true. • People like @tylercowen will blame the users, but the results just aren't what they had hoped for. Andrej Karpathy / @karpathy : Question 5 [image]
Context & Ripple Effects
GPT-4.5 arrives after GPT-4o established OpenAI’s fast, natively multimodal flagship, but reported benchmark results complicate a simple “larger is better” release narrative. The comparison matters because coding and enterprise extraction are practical workloads where model quality must justify inference cost.
The subsequent research-preview positioning of GPT-4.5 also acknowledged that it could trail OpenAI’s reasoning models. That makes the reported gap with Deep Research and other newer reasoning systems a product-segmentation issue, not merely a leaderboard result.
First-order effects
- OpenAI gains a model that reportedly improves on GPT-4o in coding benchmarks and in Box’s single-shot enterprise extraction test, giving customers another option for quality-sensitive tasks.
- The reported shortfall against reasoning models—and commentary that its price is far higher than DeepSeek v3’s—limits GPT-4.5’s straightforward value proposition for benchmark-driven buyers.
Second-order effects
- Enterprise teams will need to evaluate GPT-4.5 by task and cost rather than treating it as an automatic GPT-4o replacement; extraction gains may support selective deployment while coding and research workloads remain contested.
- Rival providers can use reasoning-model results and lower-cost positioning to compete for customers whose workloads do not reward GPT-4.5’s larger general-purpose model.
Third-order effects
- The release points toward portfolios of specialized and routed models, where general models, reasoning models, and task-specific systems compete on cost per useful result rather than on parameter scale alone.
- If performance gains remain uneven as models grow, procurement and product design will increasingly favor workload-level evaluation over flagship-model branding.
The trend: Frontier AI competition is shifting from ever-larger single models toward task-specific performance and compute-efficient model selection.