Models like o3 and Gemini 2.5 Pro feel like “Jagged AGI”: unreliable, even at some mundane tasks, but still offering superhuman capabilities in many areas
Amid today's AI boom, it's disconcerting that we still don't know how to measure how smart, creative, or empathetic these systems are.
Context & Ripple Effects
The article extends the earlier debate over whether apparent model reasoning is better understood as “jagged intelligence”, rather than a uniform increase in general capability. It also follows a period in which Gemini was seen as broadly GPT-4-class without clearly dominating benchmark comparisons, as in this early Gemini assessment.
Its significance is practical: capability claims cannot be reduced to a single intelligence score when the same model can be exceptional in some domains and fail on routine work. That makes evaluation, task selection, and oversight central to extracting value from frontier models.
First-order effects
- Teams using o3 or Gemini 2.5 Pro must validate outputs at the task level, especially where an ordinary-looking failure can invalidate otherwise strong work.
- Model providers face pressure to demonstrate reliability and scope of competence, not merely cite aggregate benchmark gains or standout demonstrations.
Second-order effects
- Enterprise buyers are likely to favor workflows that route bounded, verifiable work to models while retaining human review for brittle steps; the relevant metric becomes usefulness relative to other frontier models, not a single capability ranking.
- Competition shifts toward evaluations that expose failure modes across reasoning, creativity, and interpersonal tasks, since existing measures do not cleanly capture the differences the article describes.
Third-order effects
- If jagged performance persists, AI adoption will be organized around specialized, auditable tasks rather than an assumption that one model can safely replace broad knowledge work.
- The industry may increasingly compete on the efficiency of converting costly model development into dependable capability, as uneven reliability limits how much raw model progress translates into deployable value.
The trend: Frontier AI is moving from headline model comparisons toward task-specific reliability measurement and workflow design.