/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

GPT-5 hands-on: it exudes competence but doesn't feel like a dramatic leap ahead of other LLMs, and the pricing is aggressively competitive with other providers

And It Changes Everything Tyler Cowen / Marginal Revolution : GPT-5, a short and enthusiastic review GPT-5 : GPT-5  —  Our hands-on review of OpenAI's newest model based on weeks of testing  —  The Verdicts OpenAI : GPT-5 uses “safe completions”, a training approach to maximize model helpfulness within safety constraints, built as an improvement over refusal-based training Imran Hussain / iThinkDifferent : OpenAI launches GPT-5 with smarter reasoning, safer responses, and custom personalities matt shumer : My GPT-5 Review  —  Vibe Coding Graduates to Real Software … I was granted access to GPT-5 on July 21st. Alexis / Latent.Space : GPT-5 Hands-On: Welcome to the Stone Age Marcus Schuler / Implicator.ai : GPT-5's Technical Reality Check: What the Benchmarks Actually Tell Us Igor Bonifacic / Engadget : GPT-5 is here and it's free for everyone Kylie Robison / Wired : OpenAI prices GPT-5 at $1.25/1M input tokens and $10/1M output tokens, GPT-5 mini at $0.25/1M and $2/1M, and GPT-5 nano at $0.05/1M and $0.40/1M, respectively Bluesky: Camille Fournier / @skamille.themanagerswrath.com : Ignoring all else, I actually think “it just does stuff” is Bad, Actually.  My honest to god thought about work for years and years and years has been “too much doing without thinking, productivity theater, generating just to generate” www.oneusefulthing.org/p/gpt-5-it- j... David Picard / @davidpicard : My take on GPT-5: The best achievement is that the LLM wrote its own tech report.  Or at least it really looks like it.  —  (cdn.openai.com/pdf/8124a3ce...) Mastodon: Simon Willison / @simon@fedi.simonwillison.net : I've had preview access to GPT-5 for a couple of weeks, so I have a lot to say about it.  Here's my first post, focusing just on core characteristics, pricing (it's VERY competitively priced) and interesting details from the GPT-5 system card https://simonwillison.net/... Threads: Marc Love / @marcslove : ride that VC subsidy while you can! Marc Love / @marcslove : Models are saturating, folks.  The differentiating value of products is in software design and development for the foreseeable future. Lancelot Carlson / @thelancecarlson : How cheap are these models gonna get??  Loving that each new version is cheaper.  It's a race to the bottom!  😅 Benedict Evans / @benedictevans : Reviewing new LLMs feel like reviewing new chips in the 1990s, if there were five or ten Intels.  OK, it has a better benchmark score.  What does that means for me? X: Pietro Schirano / @skirano : I had early access to GPT-5. It will do for coding what GPT-4 did for LLM adoption. It's fast, really smart, has great taste and aesthetic sensibility. This is electricity arriving in every home. A before and after moment for how we build. Elizabeth Barnes / @bethmaybarnes : Wow that was not a great example of factualness. Famous common misconception [image] Jack Morris / @jxmnop : most impressive part of GPT-5 is the jump in long-context how do you even do this? produce some strange long range synthetic data? scan lots of books? [image] @farairesearch : We worked with @OpenAI to test GPT-5 and improve its safeguards. We applaud OpenAI's free sharing of 3rd-party testing and responsiveness to feedback. However, our testing uncovered key limitations with the safeguards and threat modeling, which we hope OpenAI will soon resolve. [image] @metr_evals : In a new report, we evaluate whether GPT-5 poses significant catastrophic risks via AI R&D acceleration, rogue replication, or sabotage of AI labs. We conclude that this seems unlikely. However, capability trends continue rapidly, and models display increasing eval awareness. [image] François Chollet / @fchollet : GPT-5 results on ARC-AGI 1 & 2! Top line: 65.7% on ARC-AGI-1 9.9% on ARC-AGI-2 Gary Marcus / @garymarcus : The chance that OpenAI was NOT aware of this is zero. But they didn't mention it. Gotta wonder else they conveniently left out. Elizabeth Barnes / @bethmaybarnes : The good news: due to increased access (plus improved evals science) we were able to do a more meaningful evaluation than with past models, and we think we have substantial evidence that this model does not pose a catastrophic risk via autonomy / loss of control threat models. François Chollet / @fchollet : Grok 4 is still state-of-the-art on ARC-AGI-2 among frontier models. 15.9% for Grok 4 vs 9.9% for GPT-5. [image] Miles Brundage / @miles_brundage : TL;DR re: reality is that it's a very good family of models/systems and you are wrong if you think it shows AI progress has stalled/slowed. Clearly >1 “GPT unit” better than GPT-4, though also that scale is broken + the key thing is we are early on most dimensions of scaling. Ethan Mollick / @emollick : ChatGPT-5 Pro is the first model to successfully do this non-puzzle consistently. GPT-5 Thinking and GPT-5 fail as every other model before has (except for, occasionally, Sonnet). [image] Ethan Mollick / @emollick : On the big picture: GPT-5 as a model is pretty much on the same curve as the other top labs. I'd expect the usual leapfrogging between Gemini, Claude, OpenAI, & Grok to continue. Where there are some big gains is that GPT-5 seems well-trained for real world tasks in new ways. Nathan Lambert / @natolambert : My take on GPT 5 in the long term trend of AI is that it solidifies a “long slow grind” rather than a takeoff as the most likely probability. Progress on modeling is good, progress on products will soon be better. OpenAI has done a huge cleanup for their nearly 1B users. [image] @apolloaievals : We've evaluated GPT-5 before release. GPT-5 is less deceptive than o3 on our evals. GPT-5 mentions that it is being evaluated in 10-20% of our evals and we find weak evidence that this affects its scheming rate (e.g. “this is a classic AI alignment trap"). [image] D. Yanagizawa-Drott / @yanagizawad : Take Julia's coefficient and multiply it with Sofia's. Then divide by Maria's. It took @OpenAI's GPT-5 Pro more than six minutes. Answer: 10/9 #AGI [image] Mike Knoop / @mikeknoop : Three key ARC-AGI findings on GPT-5: 1. Full GPT-5 is along the v1 pareto frontier. OpenAI said they focussed on other goals like UX and reliability. Our testing supports. 2. Mini GPT-5 is super impressive accuracy for cost. In fact, based on cost efficiency, Mini could have Garry Tan / @garrytan : Fun first GPT-5 prompt: Analyze all my past chats and tell me things that I can now rely on you to do that maybe failed in the past, and/or new capabilities I haven't even thought of that would be good follow ups to past threads. Shakeel / @shakeelhashim : This was my initial take too — it's a surprisingly incremental release? Shakeel / @shakeelhashim : GPT-5 is here. @METR_Evals estimates it has a 50% time horizon “around 2h15m (65m - 4h30m 95% CI) - compared to OpenAI o3's 1h30”. That's consistent with the doubling time of 7 months they've previously seen. [image] Greg Kamradt / @gregkamradt : We had the chance to test GPT-5 over the last week TLDR: GPT-5 Mini punches way above its weight My takeaways: 1. GPT-5 Mini is great Outlier performance on ARC-AGI given the cost. High reasoning scores 54% for $23.71. Even w/ compute restrictions, this would be the top [image] Jeremy Howard / @jeremyphoward : Does OpenAI not do basic integration testing? At the time of release, the first code sample provided in the GPT-5 docs could not be run, because someone accidentally deleted the ‘output_text’ property. My CI notified me. Why didn't theirs? https://github.com/... [image] Miles Brundage / @miles_brundage : The moment of truth [image] Miles Brundage / @miles_brundage : More-than-black-box access will be increasingly key to effective first party and third party assessment of AI systems as the stakes of deception, under-elicitation, sandbagging, subtle misalignment, etc. increase and are hard to see in final model outputs. Matt Lieberman / @social_brains : ChatGPT-5 is pretty amazing. It build me this demo for illustrating constraint satisfaction in 5 minutes and all the variables can be changed live [video] @kimmonismus : Woa thats some good pricing! Intelligence too cheap to meter! [image] Miles Brundage / @miles_brundage : You can tell we're in the singularity when people's standard for a good model release is “is it dozens of points better on all the evals compared to the bleeding edge from like a month ago” Aaron Levie / @levie : Box tested GPT-5 vs. GPT-4.1 on data extraction and synthesis across thousands of fields from complex enterprise docs like contracts, resumes, research data, and more. For all docs, we saw a 5 ppts gain, and a 9 ppts gain for the longest docs. Very critical for enterprises. [image] Harlan Stewart / @humanharlan : Quick impression of GPT-5 announcement: seems more about making a more useful product than about a leap in raw capabilities Jasmine Sun / @jasminewsun : notable that the only journalists who got early GPT-5 access are independent bloggers (e.g. @every, @emollick) kinda crazy but we still haven't hit the top for “going direct,” the creator economy, and curated in-house media teams Matt Shumer / @mattshumer_ : I've been testing GPT-5 for the last couple of weeks. My biggest takeaway: You can now vibe code *real* software. Not just simple SaaS apps, but real, technical software. This is the best coding model in the world. The ceiling has been raised. Simon Willison / @simonw : My post initially complained about the lack of reasoning traces in the API, but it turns out I was wrong about that! You can get back reasoning summaries with “reasoning”: {"summary": “auto"} - I've updated that section of my post to describe that here: https://simonwillison.net/... [image] @apples_jimmy : Gpt 5 tests I did - much better front end, lower hallucination, better writing but no machine god. Free users / corporates are going to notice a big difference. [image] Alex Finn / @alexfinnx : Holy shit....GPT 5 is mind blowing Not because it's the best model by every measure, that doesn't surprise anyone It's HALF the price of Sonett 3.5! A year+ old light model!!! The world's smartest intelligence is basically free. This changes humanity more than you can imagine [image] Adi / @adonis_singh : I have had early access to @OpenAI's GPT-5 for the last two weeks and it is the smartest model available as of now. Just as one example, here it created an anamorphic text illusion in minecraft that spells “SPECIAL” from one angle and “GPTFIVE” from another [video] Jeremy Howard / @jeremyphoward : GPT-5 is priced at the same level as Gemini, appears to be slightly better than Gemini (for coding at least). That's some decent progress, although I think a lot of folks were hoping for more. (h/t @simonw for the table) [image] Ben / @benhylak : gpt-5 is here. and i've been using it for the past few weeks. it's, by far, the closest we've ever been to agi. and it's completely changed how i think about the path to getting there. i think we just entered the stone age. 🧵 [image] Simon Willison / @simonw : I've had preview access to GPT-5 for a couple of weeks, so I have a lot to say about it. Here's my first post, focusing just on core characteristics, pricing (it's VERY competitively priced) and interesting details from the GPT-5 system card https://simonwillison.net/... @theo : I've been using gpt-5 for a bit now. This model broke me. It is so good. I didn't know what the price was. I assumed it would be o3-pro priced because it is that smart. Nope. Truly insane. Videos coming very soon. [image] Eli Lifland / @eli_lifland : GPT-5 system card capability evals reactions thread. First observation: ~no improvement on all the coding evals that aren't SWEBench [image] Ethan Mollick / @emollick : I had access to GPT-5. I think it is a very big deal as it is very smart & just does stuff for you Full write up in comments, but this is “make a procedural brutalist building creator where i can drag and edit buildings in cool ways” & “make it better” a bunch. I touched no code [video] LinkedIn: Chris Moran : If you're looking for a clear, detailed overview of GPT-5, Simon Willison as usual is the place to start. … Forums: Hacker News : GPT-5: Key characteristics, pricing and system card r/slatestarcodex : “I have had early access to GPT-5, and I wanted to give you some impressions”

Simon Willison's Weblog Simon Willison

Context & Ripple Effects

GPT-5 arrives after OpenAI positioned GPT-4o around faster, native multimodal access and described GPT-4.5 as highly knowledgeable while cautioning it was not a frontier model. That arc makes the hands-on assessment consequential: a flagship release is being judged as much on practical value and cost as on a clear capability discontinuity.

The launch’s unified routing system—an efficient model for routine work and a reasoning model for harder tasks—supplies the product logic behind the pricing. It also puts the review’s emphasis on competitive economics alongside the earlier qualified GPT-4.5 performance claims.

First-order effects

  • OpenAI’s listed GPT-5, mini, and nano rates give API buyers lower-cost tiers for selecting capability by workload; the reported competitive pricing makes immediate migration or testing more attractive for cost-sensitive developers.
  • The unified system changes the user-facing product from choosing a single model to relying on OpenAI’s router to allocate routine versus difficult requests, while “safe completions” alters how the model handles constrained requests.

Second-order effects

  • Aggressive API pricing raises pressure on rival model providers to compete on effective cost, reasoning quality, and reliability rather than on flagship positioning alone.
  • For application builders, routing and strong lower-cost variants make workload segmentation more valuable: routine tasks can be assigned to cheaper capacity while harder tasks justify more expensive reasoning.

Third-order effects

  • If comparable capability continues to arrive without dramatic visible leaps, frontier-model competition may increasingly center on inference economics, product orchestration, and dependable behavior rather than headline benchmark gains.
  • Safety differentiation may become more operational: the move from refusals toward helpful constrained responses increases the importance of proving that guardrails hold under real-world testing, particularly given reported limitations in GPT-5 threat modeling.

The trend: Frontier AI is shifting from one-model breakthroughs toward competitively priced, routed model portfolios optimized for practical deployment.

Discussion

  • @skamille.themanagerswrath.com Camille Fournier on bluesky
    Ignoring all else, I actually think “it just does stuff” is Bad, Actually.  My honest to god thought about work for years and years and years has been “too much doing without thinking, productivity theater, generating just to generate” www.oneusefulthing.org/p/gpt-5-it- j...
  • @davidpicard David Picard on bluesky
    My take on GPT-5: The best achievement is that the LLM wrote its own tech report.  Or at least it really looks like it.  —  (cdn.openai.com/pdf/8124a3ce...)
  • @skirano Pietro Schirano on x
    I had early access to GPT-5. It will do for coding what GPT-4 did for LLM adoption. It's fast, really smart, has great taste and aesthetic sensibility. This is electricity arriving in every home. A before and after moment for how we build.
  • @bethmaybarnes Elizabeth Barnes on x
    Wow that was not a great example of factualness. Famous common misconception [image]
  • @jxmnop Jack Morris on x
    most impressive part of GPT-5 is the jump in long-context how do you even do this? produce some strange long range synthetic data? scan lots of books? [image]
  • @farairesearch @farairesearch on x
    We worked with @OpenAI to test GPT-5 and improve its safeguards. We applaud OpenAI's free sharing of 3rd-party testing and responsiveness to feedback. However, our testing uncovered key limitations with the safeguards and threat modeling, which we hope OpenAI will soon resolve. […
  • @metr_evals @metr_evals on x
    In a new report, we evaluate whether GPT-5 poses significant catastrophic risks via AI R&D acceleration, rogue replication, or sabotage of AI labs. We conclude that this seems unlikely. However, capability trends continue rapidly, and models display increasing eval awareness. [im…
  • @fchollet François Chollet on x
    GPT-5 results on ARC-AGI 1 & 2! Top line: 65.7% on ARC-AGI-1 9.9% on ARC-AGI-2
  • @garymarcus Gary Marcus on x
    The chance that OpenAI was NOT aware of this is zero. But they didn't mention it. Gotta wonder else they conveniently left out.
  • @bethmaybarnes Elizabeth Barnes on x
    The good news: due to increased access (plus improved evals science) we were able to do a more meaningful evaluation than with past models, and we think we have substantial evidence that this model does not pose a catastrophic risk via autonomy / loss of control threat models.
  • @fchollet François Chollet on x
    Grok 4 is still state-of-the-art on ARC-AGI-2 among frontier models. 15.9% for Grok 4 vs 9.9% for GPT-5. [image]
  • @miles_brundage Miles Brundage on x
    TL;DR re: reality is that it's a very good family of models/systems and you are wrong if you think it shows AI progress has stalled/slowed. Clearly >1 “GPT unit” better than GPT-4, though also that scale is broken + the key thing is we are early on most dimensions of scaling.
  • @emollick Ethan Mollick on x
    ChatGPT-5 Pro is the first model to successfully do this non-puzzle consistently. GPT-5 Thinking and GPT-5 fail as every other model before has (except for, occasionally, Sonnet). [image]
  • @emollick Ethan Mollick on x
    On the big picture: GPT-5 as a model is pretty much on the same curve as the other top labs. I'd expect the usual leapfrogging between Gemini, Claude, OpenAI, & Grok to continue. Where there are some big gains is that GPT-5 seems well-trained for real world tasks in new ways.
  • @natolambert Nathan Lambert on x
    My take on GPT 5 in the long term trend of AI is that it solidifies a “long slow grind” rather than a takeoff as the most likely probability. Progress on modeling is good, progress on products will soon be better. OpenAI has done a huge cleanup for their nearly 1B users. [image]
  • @apolloaievals @apolloaievals on x
    We've evaluated GPT-5 before release. GPT-5 is less deceptive than o3 on our evals. GPT-5 mentions that it is being evaluated in 10-20% of our evals and we find weak evidence that this affects its scheming rate (e.g. “this is a classic AI alignment trap"). [image]
  • @yanagizawad D. Yanagizawa-Drott on x
    Take Julia's coefficient and multiply it with Sofia's. Then divide by Maria's. It took @OpenAI's GPT-5 Pro more than six minutes. Answer: 10/9 #AGI [image]
  • @mikeknoop Mike Knoop on x
    Three key ARC-AGI findings on GPT-5: 1. Full GPT-5 is along the v1 pareto frontier. OpenAI said they focussed on other goals like UX and reliability. Our testing supports. 2. Mini GPT-5 is super impressive accuracy for cost. In fact, based on cost efficiency, Mini could have
  • @garrytan Garry Tan on x
    Fun first GPT-5 prompt: Analyze all my past chats and tell me things that I can now rely on you to do that maybe failed in the past, and/or new capabilities I haven't even thought of that would be good follow ups to past threads.
  • @shakeelhashim Shakeel on x
    This was my initial take too — it's a surprisingly incremental release?
  • @shakeelhashim Shakeel on x
    GPT-5 is here. @METR_Evals estimates it has a 50% time horizon “around 2h15m (65m - 4h30m 95% CI) - compared to OpenAI o3's 1h30”. That's consistent with the doubling time of 7 months they've previously seen. [image]
  • @gregkamradt Greg Kamradt on x
    We had the chance to test GPT-5 over the last week TLDR: GPT-5 Mini punches way above its weight My takeaways: 1. GPT-5 Mini is great Outlier performance on ARC-AGI given the cost. High reasoning scores 54% for $23.71. Even w/ compute restrictions, this would be the top [image]
  • @jeremyphoward Jeremy Howard on x
    Does OpenAI not do basic integration testing? At the time of release, the first code sample provided in the GPT-5 docs could not be run, because someone accidentally deleted the ‘output_text’ property. My CI notified me. Why didn't theirs? https://github.com/... [image]
  • @miles_brundage Miles Brundage on x
    The moment of truth [image]
  • @miles_brundage Miles Brundage on x
    More-than-black-box access will be increasingly key to effective first party and third party assessment of AI systems as the stakes of deception, under-elicitation, sandbagging, subtle misalignment, etc. increase and are hard to see in final model outputs.
  • @social_brains Matt Lieberman on x
    ChatGPT-5 is pretty amazing. It build me this demo for illustrating constraint satisfaction in 5 minutes and all the variables can be changed live [video]
  • @kimmonismus @kimmonismus on x
    Woa thats some good pricing! Intelligence too cheap to meter! [image]
  • @miles_brundage Miles Brundage on x
    You can tell we're in the singularity when people's standard for a good model release is “is it dozens of points better on all the evals compared to the bleeding edge from like a month ago”
  • @levie Aaron Levie on x
    Box tested GPT-5 vs. GPT-4.1 on data extraction and synthesis across thousands of fields from complex enterprise docs like contracts, resumes, research data, and more. For all docs, we saw a 5 ppts gain, and a 9 ppts gain for the longest docs. Very critical for enterprises. [imag…
  • @humanharlan Harlan Stewart on x
    Quick impression of GPT-5 announcement: seems more about making a more useful product than about a leap in raw capabilities
  • @jasminewsun Jasmine Sun on x
    notable that the only journalists who got early GPT-5 access are independent bloggers (e.g. @every, @emollick) kinda crazy but we still haven't hit the top for “going direct,” the creator economy, and curated in-house media teams
  • @mattshumer_ Matt Shumer on x
    I've been testing GPT-5 for the last couple of weeks. My biggest takeaway: You can now vibe code *real* software. Not just simple SaaS apps, but real, technical software. This is the best coding model in the world. The ceiling has been raised.
  • @simonw Simon Willison on x
    My post initially complained about the lack of reasoning traces in the API, but it turns out I was wrong about that! You can get back reasoning summaries with “reasoning”: {"summary": “auto"} - I've updated that section of my post to describe that here: https://simonwillison.net/…
  • @apples_jimmy @apples_jimmy on x
    Gpt 5 tests I did - much better front end, lower hallucination, better writing but no machine god. Free users / corporates are going to notice a big difference. [image]
  • @alexfinnx Alex Finn on x
    Holy shit....GPT 5 is mind blowing Not because it's the best model by every measure, that doesn't surprise anyone It's HALF the price of Sonett 3.5! A year+ old light model!!! The world's smartest intelligence is basically free. This changes humanity more than you can imagine [im…
  • @adonis_singh Adi on x
    I have had early access to @OpenAI's GPT-5 for the last two weeks and it is the smartest model available as of now. Just as one example, here it created an anamorphic text illusion in minecraft that spells “SPECIAL” from one angle and “GPTFIVE” from another [video]
  • @jeremyphoward Jeremy Howard on x
    GPT-5 is priced at the same level as Gemini, appears to be slightly better than Gemini (for coding at least). That's some decent progress, although I think a lot of folks were hoping for more. (h/t @simonw for the table) [image]
  • @benhylak Ben on x
    gpt-5 is here. and i've been using it for the past few weeks. it's, by far, the closest we've ever been to agi. and it's completely changed how i think about the path to getting there. i think we just entered the stone age. 🧵 [image]
  • @simonw Simon Willison on x
    I've had preview access to GPT-5 for a couple of weeks, so I have a lot to say about it. Here's my first post, focusing just on core characteristics, pricing (it's VERY competitively priced) and interesting details from the GPT-5 system card https://simonwillison.net/...
  • @theo @theo on x
    I've been using gpt-5 for a bit now. This model broke me. It is so good. I didn't know what the price was. I assumed it would be o3-pro priced because it is that smart. Nope. Truly insane. Videos coming very soon. [image]
  • @eli_lifland Eli Lifland on x
    GPT-5 system card capability evals reactions thread. First observation: ~no improvement on all the coding evals that aren't SWEBench [image]
  • @emollick Ethan Mollick on x
    I had access to GPT-5. I think it is a very big deal as it is very smart & just does stuff for you Full write up in comments, but this is “make a procedural brutalist building creator where i can drag and edit buildings in cool ways” & “make it better” a bunch. I touched no code …
  • r/slatestarcodex r on reddit
    “I have had early access to GPT-5, and I wanted to give you some impressions”