/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

GPT-4o mini costs $0.15 per 1M input tokens and $0.60 per 1M output tokens, prices lower than those of Claude 3 Haiku and Gemini 1.5 Flash

GPT-4o mini.  I've been complaining about how under-powered GPT 3.5 is for the price for a while now (I made fun of it in a keynote a few weeks ago).

Simon Willison's Weblog Simon Willison

Context & Ripple Effects

This pricing move accompanied GPT-4o mini’s launch as a smaller GPT-4o variant replacing GPT-3.5 Turbo, making the comparison more consequential than a standalone price cut: OpenAI was resetting the entry-level model it offered developers.

Later coverage shows OpenAI continued to segment its lineup by capability and price, including o3-mini’s lower price relative to GPT-4o and o1. The contrast underscores that low-cost models and higher-priced reasoning models can serve distinct workloads rather than form one simple price ladder.

First-order effects

  • Developers processing high token volumes can obtain GPT-4o mini input and output at rates below the cited Claude 3 Haiku and Gemini 1.5 Flash prices, improving the immediate economics of choosing OpenAI for eligible workloads.
  • Replacing GPT-3.5 Turbo with this lower-cost model shifts OpenAI’s entry-level API offering toward a newer model family and gives existing budget-sensitive users a migration target.

Second-order effects

  • Anthropic and Google face a clearer price-performance benchmark in the low-cost model tier; they may need to compete through price, capability, or product differentiation rather than treating inexpensive models as a secondary segment.
  • Lower token charges can make larger-volume use cases more viable, but buyers will increasingly compare effective cost for a completed task—not token price alone—across providers.

Third-order effects

  • If comparable capability continues to move into cheaper models, the market may separate into commoditized high-volume inference and premium tiers for harder workloads, with providers differentiating through reliability, modalities, and specialized performance.
  • The later appearance of both cheaper mini models and materially pricier frontier offerings suggests model pricing will remain segmented rather than decline uniformly; the balance depends on whether lower-cost models meet production needs.

The trend: This is one data point in the inference-cost curve: vendors are using lower-priced smaller models to broaden API adoption while preserving premium pricing for more demanding capability tiers.

Discussion

  • @sama Sam Altman on x
    way back in 2022, the best model in the world was text-davinci-003. it was much, much worse than this new model. it cost 100x more.
  • @abacaj Anton on x
    openai just effectively made llama 3 70B obsolete (outside of being able to fine tune it for cheaper) [image]
  • @nickadobos Nick Dobos on x
    Anthropic: “Yeah check it out sonnet 3.5 is pretty good, state of the art” OpenAI: “That's a cool profit margin you have there. Would be a shame if we cut your prices in half”
  • @elidourado Eli Dourado on x
    Just occurred to me to run these numbers. GPT-4o is 87 tokens per second and $15 per million output tokens, so that works out to a wage of $4.70 per hour. GPT-4o mini: 183 tps @ $0.60 per MTok = $0.39/hour. A single instance outputting tokens all day would be under $10.
  • @jam3scampbell James Campbell on x
    I added gpt-4o-mini to Proctor and it's insane how cheap it is Put it this way: if you send a screenshot into GPT-4o-mini every minute of every day 24/7 for an entire month, the cost comes to $7. That's 24/7 non-stop *intelligent* video surveillance for the cost of a Netflix [ima…
  • @mehedih_ Mehedi on x
    i thought i was going crazy, but 4o mini does indeed have the same $$ as 4o. the crazier thing is openai's blog post doesn't mention it at all 🙃
  • @sandersted Ted Sanders on x
    One reason you don't see AI everywhere yet is that marginal costs are high. Hard to run a game or site that burns, say, $0.25 / user / day. But costs are plummeting. Today's gpt-4o mini is ~99% cheaper than original gpt-3.5, at like ~30 pages a penny. https://openai.com/...
  • @xlr8harder @xlr8harder on x
    I was not seeing this in my feed, so here's the price difference between haiku and gpt4o mini. gpt4o mini - $0.15 / 1M input, $0.60 / 1M output haiku - $0.25 / 1M input, $1.25 / 1M output Interested in subjective takes from people who heavily use Haiku.
  • @_louiepeters Louie Peters on x
    To continue this trend and evidence of rapid AI/ LLM deflation; OpenAI's new GPT-4o-mini out today is an average ~140x cheaper than GPT-4 was at release in March 2023 (also mostly better!), while ~230x cheaper and vastly better than Da-Vinci 002 in Aug-22 (the best model at the […
  • @simonw Simon Willison on x
    Also notable: the model has 128,000 input tokens and 16,000 output tokens, twice the output tokens of Claude 3.5 Sonnet, making it a better fit for translation and transformation tasks where output length more closely matches the input
  • @mckaywrigley Mckay Wrigley on x
    1 million tokens from the *best* AI model costs like $3. Do you understand what a miracle that is? For the price of a McDouble I can have the world's best LLM write me 20k lines of code. Pure magic.
  • @simonw Simon Willison on x
    My notes on today's release of GPT-4o mini: https://simonwillison.net/... The biggest news is the price: this is cheaper even than Claude 3 Haiku, at just 15c per million input tokens / 60c per million output, a 99% reduction on the prices of GPT-3 Da Vinci from 2022!
  • @openaidevs @openaidevs on x
    Introducing GPT-4o mini! It's our most intelligent and affordable small model, available today in the API. GPT-4o mini is significantly smarter and cheaper than GPT-3.5 Turbo. https://openai.com/... [image]
  • @sama Sam Altman on x
    towards intelligence too cheap to meter: https://openai.com/... 15 cents per million input tokens, 60 cents per million output tokens, MMLU of 82%, and fast. most importantly, we think people will really, really like using the new model.
  • @openai @openai on x
    We're continuing to make advanced AI accessible to all with the launch of GPT-4o mini, now available in the API and rolling out in ChatGPT today.
  • @rasbt Sebastian Raschka on x
    @karpathy A positive, refreshing example of this was Gemma-2's knowledge distillation from the 27B model into the smaller versions (and also MiniLM before that). However, it's also true that MMLU is just multiple-choice answering. It's a good indicator of knowledge, but as we all…
  • @karpathy Andrej Karpathy on x
    This is not very different from Tesla with self-driving networks. What is the “offline tracker” (presented in AI day)? It is a synthetic data generating process, taking the previous, weaker (or e.g. singleframe, or bounding box only) models, running them over clips in an offline
  • @tobi Tobi Lutke on x
    16k output tokens is an significant update
  • @luke_metro @luke_metro on x
    Guys I don't think AI is going to destroy our power grid
  • @artificialanlys @artificialanlys on x
    GPT-4o Mini, announced today, is very impressive for how cheap it is being offered 👀 With a MMLU score of 82% (reported by TechCrunch), it surpasses the quality of other smaller models including Gemini 1.5 Flash (79%) and Claude 3 Haiku (75%). What is particularly exciting is [im…
  • @elonmusk Elon Musk on x
    @karpathy Yup, same thing happening with real-world AI at Tesla
  • @bilawalsidhu Bilawal Sidhu on x
    Make big models, make good data, make models smoler It's like fractional distillation of data instead of oil It's a staircase — each step more refined Aiming for the ‘perfect training set’ - the ultimate concentrate of all human knowledge and creativity Whether Gemini 1.5
  • @jeffintime Jeff Harris on x
    hidden gem: GPT-4o mini supports 16K max_tokens (up from 4K on GPT-4T+GPT-4o) https://openai.com/... [image]
  • @romainhuet Romain Huet on x
    Such an incredible time to be a developer! GPT-4o mini keeps pushing the cost-intelligence frontier—the cost per token has dropped by 99% since 2022. Model launches at @OpenAI are always special, and this one happens to be on my birthday! Thanks for the gift of GPT-4o mini! 🎉
  • @emollick Ethan Mollick on x
    First impressions with GPT-4o-mini (what a name) is that it is impressive for a small model but no replacement for a frontier model. When given complex education prompts it can't follow instructions as well & misses nuance GPT-4o nails I hope it will not be the only free model. […
  • @garymarcus Gary Marcus on x
    OpenAI's price cuts are right in line with predictions #2, #3, #4, and #7 below; all seven still look to be on track. As I warned last August, generative AI may turn out to be a dud.
  • @miramurati Mira Murati on x
    GPT-4o mini makes intelligence far more affordable opening up a wide range of applications. Available on API and rolling out on ChatGPT today
  • @karpathy Andrej Karpathy on x
    LLM model size competition is intensifying… backwards!  My bet is that we'll see models that “think” very well and reliably that are very very small.  There is most likely a setting even of GPT-2 parameters for which most people will consider GPT-2 “smart”.  The reason current …
  • @gdb Greg Brockman on x
    We built gpt-4o mini due to popular demand from developers. We ❤️ developers, and aim to provide them the best tools to convert machine intelligence into positive applications across every domain. Please keep the feedback coming.
  • @romainhuet Romain Huet on x
    Launching GPT-4o mini: the most intelligent and cost-efficient small model yet! Smarter and cheaper than GPT-3.5 Turbo, it's ideal for function calling, large contexts, real-time interactions—and has vision capabilities. Can't wait to see what you build! https://openai.com/...
  • @andrewcurran_ Andrew Curran on x
    For the many people asking for the API pricing for GPT-4o Mini it is: 15¢ per M token input 60¢ per M token output 128k context window
  • @andrewcurran_ Andrew Curran on x
    According to The Verge MMLU benchmarks look like this: GPT-4o - 88.7 GPT-4o Mini - 82 GPT-3.5- 70 https://x.com/... [image]
  • @andrewcurran_ Andrew Curran on x
    API pricing from TechCrunch. 15 cents/M input 60 cents/M output Context length 128k [image]
  • @andrewcurran_ Andrew Curran on x
    From Bloomberg: 'OpenAI also said that GPT-4o mini is the company's first AI model to use a new safety tactic it developed called “instruction hierarchy.” [image]
  • @apples_jimmy @apples_jimmy on x
    3.5 getting the boot, finally.
  • @mattshumer_ Matt Shumer on x
    New @OpenAI model! GPT-4o mini drops today. Seems to be a replacement for GPT-3.5-Turbo (finally!) Seems like this model will be very similar to Claude Haiku — fast / cheap, and very good at handling few-shot prompts [image]
  • @rachelmetz Rachel Metz on x
    New day, new model; OpenAI is rolling out GPT-4o mini, which will replace GPT-3.5 in ChatGPT. OpenAI Releases GPT-4o Mini, a Cheaper Version of Flagship AI Model https://www.bloomberg.com/...
  • @andrewcurran_ Andrew Curran on x
    Here's our new model. ‘GPT-4o mini’. ‘the most capable and cost-efficient small model available today’ according to OpenAI. Going live today for free and pro. [image]
  • r/artificial r on reddit
    One-Minute Daily AI News 7/18/2024
  • r/OpenAI r on reddit
    GPT-4o mini: advancing cost-efficient intelligence
  • r/LocalLLaMA r on reddit
    New OPENAI Model
  • r/ChatGPT r on reddit
    OpenAI Slashes the Cost of Using Its AI With a “Mini” Model
  • r/singularity r on reddit
    OpenAI debuts mini version of its most powerful model yet