/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

Models like o3 and Gemini 2.5 Pro feel like “Jagged AGI”: unreliable, even at some mundane tasks, but still offering superhuman capabilities in many areas

Amid today's AI boom, it's disconcerting that we still don't know how to measure how smart, creative, or empathetic these systems are.

One Useful Thing Ethan Mollick

Context & Ripple Effects

The article extends the earlier debate over whether apparent model reasoning is better understood as “jagged intelligence”, rather than a uniform increase in general capability. It also follows a period in which Gemini was seen as broadly GPT-4-class without clearly dominating benchmark comparisons, as in this early Gemini assessment.

Its significance is practical: capability claims cannot be reduced to a single intelligence score when the same model can be exceptional in some domains and fail on routine work. That makes evaluation, task selection, and oversight central to extracting value from frontier models.

First-order effects

  • Teams using o3 or Gemini 2.5 Pro must validate outputs at the task level, especially where an ordinary-looking failure can invalidate otherwise strong work.
  • Model providers face pressure to demonstrate reliability and scope of competence, not merely cite aggregate benchmark gains or standout demonstrations.

Second-order effects

  • Enterprise buyers are likely to favor workflows that route bounded, verifiable work to models while retaining human review for brittle steps; the relevant metric becomes usefulness relative to other frontier models, not a single capability ranking.
  • Competition shifts toward evaluations that expose failure modes across reasoning, creativity, and interpersonal tasks, since existing measures do not cleanly capture the differences the article describes.

Third-order effects

  • If jagged performance persists, AI adoption will be organized around specialized, auditable tasks rather than an assumption that one model can safely replace broad knowledge work.
  • The industry may increasingly compete on the efficiency of converting costly model development into dependable capability, as uneven reliability limits how much raw model progress translates into deployable value.

The trend: Frontier AI is moving from headline model comparisons toward task-specific reliability measurement and workflow design.

Discussion

  • @rizzn Mark Rizzn Hopkins on x
    @emollick o3 is ultra tuned but it can barely hold a conversation... it's very in its own world. Smart? Yes. Social skills? Almost non existent.
  • r/artificial r on reddit
    On Jagged AGI: o3, Gemini 2.5, and everything after