GPT-4o mini costs $0.15 per 1M input tokens and $0.60 per 1M output tokens, prices lower than those of Claude 3 Haiku and Gemini 1.5 Flash
GPT-4o mini. I've been complaining about how under-powered GPT 3.5 is for the price for a while now (I made fun of it in a keynote a few weeks ago).
Simon Willison's Weblog Simon Willison
Context & Ripple Effects
This pricing move accompanied GPT-4o mini’s launch as a smaller GPT-4o variant replacing GPT-3.5 Turbo, making the comparison more consequential than a standalone price cut: OpenAI was resetting the entry-level model it offered developers.
Later coverage shows OpenAI continued to segment its lineup by capability and price, including o3-mini’s lower price relative to GPT-4o and o1. The contrast underscores that low-cost models and higher-priced reasoning models can serve distinct workloads rather than form one simple price ladder.
First-order effects
- Developers processing high token volumes can obtain GPT-4o mini input and output at rates below the cited Claude 3 Haiku and Gemini 1.5 Flash prices, improving the immediate economics of choosing OpenAI for eligible workloads.
- Replacing GPT-3.5 Turbo with this lower-cost model shifts OpenAI’s entry-level API offering toward a newer model family and gives existing budget-sensitive users a migration target.
Second-order effects
- Anthropic and Google face a clearer price-performance benchmark in the low-cost model tier; they may need to compete through price, capability, or product differentiation rather than treating inexpensive models as a secondary segment.
- Lower token charges can make larger-volume use cases more viable, but buyers will increasingly compare effective cost for a completed task—not token price alone—across providers.
Third-order effects
- If comparable capability continues to move into cheaper models, the market may separate into commoditized high-volume inference and premium tiers for harder workloads, with providers differentiating through reliability, modalities, and specialized performance.
- The later appearance of both cheaper mini models and materially pricier frontier offerings suggests model pricing will remain segmented rather than decline uniformly; the balance depends on whether lower-cost models meet production needs.
The trend: This is one data point in the inference-cost curve: vendors are using lower-priced smaller models to broaden API adoption while preserving premium pricing for more demanding capability tiers.
Related: Effective inference cost · AI cost per useful task · Inference Cost Curve · GPT-4o mini launch and GPT-3.5 Turbo replacement · OpenAI o3-mini pricing · Gemini
Related Coverage
- Daily Digest: Mini models are insane Ben's Bites
- OpenAI launches GPT-4o Mini: A cheaper, lighter AI model for developers Business Today · Pranav Dixit
- GPT-4o mini: advancing cost-efficient intelligence OpenAI
- Launching GPT-4o mini: the most intelligent and cost-efficient small model yet! This is a huge step in making intelligence as broadly accessible as possible … Romain Huet
- OpenAI is releasing a cheaper, smarter model The Verge · Kylie Robison
- OpenAI Releases GPT-4o Mini, a Cheaper Version of Flagship AI Model Bloomberg · Rachel Metz
- OpenAI reveals cheaper mini version of its flagship GPT-4o Silicon Republic · Leigh Mc Gowran
- OpenAI Slashes the Cost of Using Its AI With a “Mini” Model Wired · Will Knight
- OpenAI's 4o-mini brings big brains on a budget. Ben's Bites
- OpenAI unveils GPT-4o mini, a smaller and cheaper AI model TechCrunch · Maxwell Zeff
- OpenAI brings GPT-4o mini AI model targeting app developers: Check details Business Standard · Harsh Shivam
- OpenAI's unveils GPT-4o Mini! The Rundown AI · Rowan Cheung
- Here's the real reason AI companies are slimming down their models Fast Company · Mark Sullivan
- OpenAI launches GPT-4o mini, which will replace GPT-3.5 in ChatGPT Ars Technica · Benj Edwards
- OpenAI Introduces GPT-4o Mini: A More Efficient & Cost-Effective AI Model Blockonomi · Oliver Dale
- Microsoft announces safety and performance upgrades for Azure OpenAI Service Neowin · Pradeep Viswanathan
- OpenAI Introduced A New Small AI Model, ChatGPT-4o mini That Is Faster, Cost Efficient, And Could Outperform Others Wccftech · Ezza Ijaz
- OpenAI's new GPT-4o mini slashes costs and tackles text, code, and beyond Tech Funding News · Vivek Chhetri
- OpenAI goes lite, releases cheaper AI model called GPT-4o mini Cryptopolitan · Jeffrey Gogo
- OpenAI unveils cheaper small AI model GPT-4o mini Reuters · Deborah Sophia
- OpenAI offers GPT-4o mini to slash the cost of applications ZDNet · Tiernan Ray
- We hear all the time from our customers about the need for faster, cheaper models. — So very excited for the launch of GPT-4o mini, second in our “omni” model family. … Amol Shah
- Today, we're announcing GPT-4o mini, our most cost-efficient small model. We expect GPT-4o mini will significantly expand the range … Jenny O'Leary
- OpenAI just launched a new model: We're announcing #GPT-4o mini, our most cost-efficient small model. … Matt Kroll
- Today, OpenAI introduced GPT4o-mini, an affordable, fast, and smart model. — 🚀 Already available for testing on #AzureOpenAI (early access playground) … Mick Vleeshouwer
Discussion
-
@sama
Sam Altman
on x
way back in 2022, the best model in the world was text-davinci-003. it was much, much worse than this new model. it cost 100x more.
-
@abacaj
Anton
on x
openai just effectively made llama 3 70B obsolete (outside of being able to fine tune it for cheaper) [image]
-
@nickadobos
Nick Dobos
on x
Anthropic: “Yeah check it out sonnet 3.5 is pretty good, state of the art” OpenAI: “That's a cool profit margin you have there. Would be a shame if we cut your prices in half”
-
@elidourado
Eli Dourado
on x
Just occurred to me to run these numbers. GPT-4o is 87 tokens per second and $15 per million output tokens, so that works out to a wage of $4.70 per hour. GPT-4o mini: 183 tps @ $0.60 per MTok = $0.39/hour. A single instance outputting tokens all day would be under $10.
-
@jam3scampbell
James Campbell
on x
I added gpt-4o-mini to Proctor and it's insane how cheap it is Put it this way: if you send a screenshot into GPT-4o-mini every minute of every day 24/7 for an entire month, the cost comes to $7. That's 24/7 non-stop *intelligent* video surveillance for the cost of a Netflix [ima…
-
@mehedih_
Mehedi
on x
i thought i was going crazy, but 4o mini does indeed have the same $$ as 4o. the crazier thing is openai's blog post doesn't mention it at all 🙃
-
@sandersted
Ted Sanders
on x
One reason you don't see AI everywhere yet is that marginal costs are high. Hard to run a game or site that burns, say, $0.25 / user / day. But costs are plummeting. Today's gpt-4o mini is ~99% cheaper than original gpt-3.5, at like ~30 pages a penny. https://openai.com/...
-
@xlr8harder
@xlr8harder
on x
I was not seeing this in my feed, so here's the price difference between haiku and gpt4o mini. gpt4o mini - $0.15 / 1M input, $0.60 / 1M output haiku - $0.25 / 1M input, $1.25 / 1M output Interested in subjective takes from people who heavily use Haiku.
-
@_louiepeters
Louie Peters
on x
To continue this trend and evidence of rapid AI/ LLM deflation; OpenAI's new GPT-4o-mini out today is an average ~140x cheaper than GPT-4 was at release in March 2023 (also mostly better!), while ~230x cheaper and vastly better than Da-Vinci 002 in Aug-22 (the best model at the […
-
@simonw
Simon Willison
on x
Also notable: the model has 128,000 input tokens and 16,000 output tokens, twice the output tokens of Claude 3.5 Sonnet, making it a better fit for translation and transformation tasks where output length more closely matches the input
-
@mckaywrigley
Mckay Wrigley
on x
1 million tokens from the *best* AI model costs like $3. Do you understand what a miracle that is? For the price of a McDouble I can have the world's best LLM write me 20k lines of code. Pure magic.
-
@simonw
Simon Willison
on x
My notes on today's release of GPT-4o mini: https://simonwillison.net/... The biggest news is the price: this is cheaper even than Claude 3 Haiku, at just 15c per million input tokens / 60c per million output, a 99% reduction on the prices of GPT-3 Da Vinci from 2022!
-
@openaidevs
@openaidevs
on x
Introducing GPT-4o mini! It's our most intelligent and affordable small model, available today in the API. GPT-4o mini is significantly smarter and cheaper than GPT-3.5 Turbo. https://openai.com/... [image]
-
@sama
Sam Altman
on x
towards intelligence too cheap to meter: https://openai.com/... 15 cents per million input tokens, 60 cents per million output tokens, MMLU of 82%, and fast. most importantly, we think people will really, really like using the new model.
-
@openai
@openai
on x
We're continuing to make advanced AI accessible to all with the launch of GPT-4o mini, now available in the API and rolling out in ChatGPT today.
-
@rasbt
Sebastian Raschka
on x
@karpathy A positive, refreshing example of this was Gemma-2's knowledge distillation from the 27B model into the smaller versions (and also MiniLM before that). However, it's also true that MMLU is just multiple-choice answering. It's a good indicator of knowledge, but as we all…
-
@karpathy
Andrej Karpathy
on x
This is not very different from Tesla with self-driving networks. What is the “offline tracker” (presented in AI day)? It is a synthetic data generating process, taking the previous, weaker (or e.g. singleframe, or bounding box only) models, running them over clips in an offline
-
@tobi
Tobi Lutke
on x
16k output tokens is an significant update
-
@luke_metro
@luke_metro
on x
Guys I don't think AI is going to destroy our power grid
-
@artificialanlys
@artificialanlys
on x
GPT-4o Mini, announced today, is very impressive for how cheap it is being offered 👀 With a MMLU score of 82% (reported by TechCrunch), it surpasses the quality of other smaller models including Gemini 1.5 Flash (79%) and Claude 3 Haiku (75%). What is particularly exciting is [im…
-
@elonmusk
Elon Musk
on x
@karpathy Yup, same thing happening with real-world AI at Tesla
-
@bilawalsidhu
Bilawal Sidhu
on x
Make big models, make good data, make models smoler It's like fractional distillation of data instead of oil It's a staircase — each step more refined Aiming for the ‘perfect training set’ - the ultimate concentrate of all human knowledge and creativity Whether Gemini 1.5
-
@jeffintime
Jeff Harris
on x
hidden gem: GPT-4o mini supports 16K max_tokens (up from 4K on GPT-4T+GPT-4o) https://openai.com/... [image]
-
@romainhuet
Romain Huet
on x
Such an incredible time to be a developer! GPT-4o mini keeps pushing the cost-intelligence frontier—the cost per token has dropped by 99% since 2022. Model launches at @OpenAI are always special, and this one happens to be on my birthday! Thanks for the gift of GPT-4o mini! 🎉
-
@emollick
Ethan Mollick
on x
First impressions with GPT-4o-mini (what a name) is that it is impressive for a small model but no replacement for a frontier model. When given complex education prompts it can't follow instructions as well & misses nuance GPT-4o nails I hope it will not be the only free model. […
-
@garymarcus
Gary Marcus
on x
OpenAI's price cuts are right in line with predictions #2, #3, #4, and #7 below; all seven still look to be on track. As I warned last August, generative AI may turn out to be a dud.
-
@miramurati
Mira Murati
on x
GPT-4o mini makes intelligence far more affordable opening up a wide range of applications. Available on API and rolling out on ChatGPT today
-
@karpathy
Andrej Karpathy
on x
LLM model size competition is intensifying… backwards! My bet is that we'll see models that “think” very well and reliably that are very very small. There is most likely a setting even of GPT-2 parameters for which most people will consider GPT-2 “smart”. The reason current …
-
@gdb
Greg Brockman
on x
We built gpt-4o mini due to popular demand from developers. We ❤️ developers, and aim to provide them the best tools to convert machine intelligence into positive applications across every domain. Please keep the feedback coming.
-
@romainhuet
Romain Huet
on x
Launching GPT-4o mini: the most intelligent and cost-efficient small model yet! Smarter and cheaper than GPT-3.5 Turbo, it's ideal for function calling, large contexts, real-time interactions—and has vision capabilities. Can't wait to see what you build! https://openai.com/...
-
@andrewcurran_
Andrew Curran
on x
For the many people asking for the API pricing for GPT-4o Mini it is: 15¢ per M token input 60¢ per M token output 128k context window
-
@andrewcurran_
Andrew Curran
on x
According to The Verge MMLU benchmarks look like this: GPT-4o - 88.7 GPT-4o Mini - 82 GPT-3.5- 70 https://x.com/... [image]
-
@andrewcurran_
Andrew Curran
on x
API pricing from TechCrunch. 15 cents/M input 60 cents/M output Context length 128k [image]
-
@andrewcurran_
Andrew Curran
on x
From Bloomberg: 'OpenAI also said that GPT-4o mini is the company's first AI model to use a new safety tactic it developed called “instruction hierarchy.” [image]
-
@apples_jimmy
@apples_jimmy
on x
3.5 getting the boot, finally.
-
@mattshumer_
Matt Shumer
on x
New @OpenAI model! GPT-4o mini drops today. Seems to be a replacement for GPT-3.5-Turbo (finally!) Seems like this model will be very similar to Claude Haiku — fast / cheap, and very good at handling few-shot prompts [image]
-
@rachelmetz
Rachel Metz
on x
New day, new model; OpenAI is rolling out GPT-4o mini, which will replace GPT-3.5 in ChatGPT. OpenAI Releases GPT-4o Mini, a Cheaper Version of Flagship AI Model https://www.bloomberg.com/...
-
@andrewcurran_
Andrew Curran
on x
Here's our new model. ‘GPT-4o mini’. ‘the most capable and cost-efficient small model available today’ according to OpenAI. Going live today for free and pro. [image]
-
r/artificial
r
on reddit
One-Minute Daily AI News 7/18/2024
-
r/OpenAI
r
on reddit
GPT-4o mini: advancing cost-efficient intelligence
-
r/LocalLLaMA
r
on reddit
New OPENAI Model
-
r/ChatGPT
r
on reddit
OpenAI Slashes the Cost of Using Its AI With a “Mini” Model
-
r/singularity
r
on reddit
OpenAI debuts mini version of its most powerful model yet