Sources: OpenAI is developing a new LLM, codenamed Garlic, that outperforms Gemini 3 and Claude Opus 4.5 in coding and reasoning tasks, per internal evaluations
OpenAI, which in recent weeks has appeared to fall behind Google in AI development, is fighting back with a new large language model codenamed Garlic. X: @amir X: Amir Efrati / @amir : new: OpenAI developed a new model, Garlic, that could bridge gap w/Google in pretraining fixing pretraining problems is ~crucial~ [image]
Context & Ripple Effects
The report lands against coverage of Google’s multimodal Gemini training push, which was described as a comeback strong enough to trigger a Code Red at OpenAI. Garlic is presented as OpenAI’s attempt to address pretraining weaknesses rather than simply claim another benchmark increment.
OpenAI had already highlighted GPT-5’s coding and reasoning benchmark results and signaled interest in reasoning-capable models. The reported internal comparison puts the competitive focus back on whether training improvements can translate into a durable lead.
First-order effects
- OpenAI gains a potential internal route to narrow its reported gap with Google if Garlic’s evaluations hold; the report does not establish a release timetable or independent performance validation.
- Google and Anthropic are the immediate comparison points, with Gemini 3 and Claude Opus 4.5 reportedly surpassed on coding and reasoning tasks in OpenAI’s internal testing.
Second-order effects
- A credible Garlic result would increase pressure on rival labs to demonstrate comparable coding and reasoning performance, especially where model quality is judged through evaluations rather than broad product availability.
- Developers and enterprise buyers may delay conclusions about model leadership until results are independently tested, making reproducible task performance more consequential than internal rankings alone.
Third-order effects
- The episode points to frontier-model competition increasingly turning on training-process advances—particularly pretraining quality—rather than on a single public launch or benchmark score.
- If this pattern persists, the advantage will favor labs able to sustain repeated training and evaluation cycles, reinforcing concentration among organizations with the resources to run them.
The trend: Frontier AI competition is becoming a contest over whether deeper training-stack improvements can repeatedly reset leadership in coding and reasoning.