/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

The GPT-4 barrier has finally been broken, with Gemini 1.5, Mistral Large, Claude 3 Opus, and Inflection-2.5 benchmarking near or even above OpenAI's model

Four weeks ago, GPT-4 remained the undisputed champion: consistently at the top of every key benchmark, but more importantly the clear winner in terms of “vibes”.

Simon Willison's Weblog

Context & Ripple Effects

Just weeks earlier, a hands-on assessment had found Gemini Advanced was already in GPT-4’s class, though it did not show an unambiguous benchmark advantage. This coverage expands that challenge from one Google model to a broader group spanning Google, Anthropic, Mistral and Inflection.

The significance is not that one benchmark has a new leader, but that GPT-4 is no longer presented as the sole reference point for top-tier general-purpose model performance.

First-order effects

  • Gemini 1.5, Mistral Large, Claude 3 Opus and Inflection-2.5 gain credible positioning as GPT-4-level alternatives on the benchmarks cited.
  • OpenAI’s GPT-4 loses the benefit of being the uncontested performance baseline, making comparative evaluation more relevant for prospective users.

Second-order effects

  • Model buyers can more plausibly shortlist competing providers rather than treating GPT-4 as the default technical choice; differences in pricing, access and product fit become more consequential once benchmark performance converges.
  • Competing labs are pushed to substantiate claims with broader evaluations, since clearing a single incumbent benchmark bar does not by itself establish a durable product advantage.

Third-order effects

  • If frontier-model parity persists, competition is likely to shift from a single leaderboard leader toward the cost per useful task, distribution and reliability of competing model stacks.
  • Benchmark leadership may become less decisive as a market signal, increasing the importance of task-specific testing in enterprise AI procurement.

The trend: Frontier AI is moving from a single-model performance hierarchy toward a multi-vendor market in which comparable capability shifts differentiation to deployment economics and product execution.

Discussion

  • @reidhoffman Reid Hoffman on threads
    An exciting milestone for Inflection and Pi users: Our latest model, Inflection-2.5, puts Pi up there with GPT-4 on all benchmarks.  And yes, Pi is still curious and kind.
  • @thesampadilla Sam Padilla on x
    Here's a contrarian take: The returns of “winning” the foundational model race are overvalued. I think domain-specific models will turn out to be preferred for high value tasks, and that is where the asymmetry and real value will be. The rest is a race for the bottom.
  • @balajis Balaji on x
    Suddenly, GPT4 has four competitors. - Gemini 1.5 - Mistral Large - Claude 3 Opus - Inflection-2.5 Links below. All are impressive. I like Claude the best. https://simonwillison.net/... [image]
  • @balajis Balaji on x
    1) Gemini 1.5 is incredible for apolitical code. Has a waitlist, though. https://blog.google/... 2) Mistral releases great open models and their Large is more suitable for devs. https://mistral.ai/... 3) Claude is the most similar to ChatGPT in that you can just sign up and... [i…
  • @garymarcus Gary Marcus on x
    5 models converging on roughly the same place. None leaping forward. Plateau? Maybe if there is only one internet and no robust solution to reasoning or factuality or compositionality or the problem of generalizing beyond the distribution, that's what you get?