OpenAI announces an 80% price drop for its o3 model and a “flex” mode for synchronous processing that charges $5 for input and $20 for output per million tokens
just cheaper. https://platform.openai.com/ ... [image] Kevin Weil / @kevinweil : Because you all asked: we're going to double the rate limits for o3 for Plus users. Rolling out as we speak. Now go do awesome stuff with it! Edwin / @edwinarbus : o3 is 20% cheaper than GPT-4o. Rethink everything. [image] Ashwin Sinha / @ashwinning : o3 is now priced in line with Gemini 2.5 Pro [image] @dkundel : The experience of using the OpenAI Agents SDK for all you Cloudflare Workers developers should be much better in the latest release. Right on time to make use of that 80% price cut of o3! Thank you @threepointone for the help! [image] Nathan Lambert / @natolambert : we love model competition! @therealadamg : @pli_cachete It's not distilled. Same model. Simon Willison / @simonw : o3 80% price drop is a big shake-up in terms of LLM pricing It's now the same as GPT 4.1 ($2/$8), less than Claude Sonnet 4 ($3/$15) and Opus 4 ($15/$75) and sits between Gemini 2.5 Pro for >200,00 tokens ($2.50/$15) and 2.5 Pro for <200,000 ($1.25/$10) https://simonwillison.net/... Aaron Levie / @levie : OpenAI dropped o3 prices by 80%. The amazing thing about AI is that use cases that are cost prohibitive today become affordable within a year. That means it's better to build apps that are super powerful and not worry about costs, than apps that are affordable but simple. Forums: Hacker News : OpenAI dropped the price of o3 by 80%
Context & Ripple Effects
OpenAI introduced the o3 family as reasoning models designed to think before responding, then established a lower-cost tier with o3-mini’s markedly cheaper token pricing. The new o3 pricing brings the flagship closer to the price bands observers cited for Gemini and GPT-4o.
The cut arrives alongside o3-pro’s higher-priced API and ChatGPT rollout, reinforcing a clearer split between an accessible general reasoning model and a premium option positioned for more demanding tool use.
First-order effects
- API customers can use o3 at an 80% lower price, while the new synchronous flex mode is priced at $5 per million input tokens and $20 per million output tokens.
- Plus users receive doubled o3 rate limits, increasing immediate access for subscribers; developers building on Cloudflare Workers can lower the cost of o3-backed workloads.
Second-order effects
- The new price point raises pressure on Gemini, Claude Sonnet 4 and other frontier-model providers whose listed prices were higher or nearby, particularly for workloads where token cost is a primary buying criterion.
- OpenAI’s lineup becomes more explicitly segmented: lower-priced o3 can serve broader production use while o3-pro’s context-heavy, tool-oriented positioning supports a premium tier.
Third-order effects
- If providers keep pairing frontier capability with steep price cuts, model selection will increasingly turn on cost per useful task rather than a model’s headline intelligence alone.
- Flex-style offers point toward capacity-aware inference pricing, where customers trade processing conditions or service characteristics against unit cost and providers use pricing to route demand across available capacity.
The trend: Frontier-model vendors are turning inference pricing and capacity tiers into core competitive levers as reasoning models move from premium experiments toward wider production deployment.