OpenAI says its Jalapeño chip delivered 1.5x-1.9x more AI work per watt and 1.7x-3.6x lower latency than Nvidia chips across GPT-OSS, DeepSeek R1, Kimi K2.5 1T
Jalapeño outperformed Nvidia's superchips on an AI inference benchmark test.
The Verge Emma Roth
Context & Ripple Effects
When OpenAI and Broadcom first unveiled Jalapeño in June, the claim was process speed — an LLM-optimized inference chip taken from design to manufacturing tape-out in about nine months. Today's benchmarks are the payoff: across GPT-OSS, DeepSeek R1, and Kimi K2.5 1T, OpenAI reports 1.5x-1.9x more work per watt and 1.7x-3.6x lower latency than Nvidia's superchips, and a SemiAnalysis deep-dive published alongside adds that the ASIC also beat AMD and Google parts on several top open-weight models.
The timing sharpens the rivalry: Nvidia spent the summer courting OpenAI as a flagship customer — Jensen Huang named it among the first big users of the new Vera CPUs in June — while separately confirming a $20 billion bet on Groq's LPU technology and putting its Groq 3 LPX inference accelerator into full production. OpenAI is answering not with a purchase but with silicon of its own.
First-order effects
- OpenAI says it will begin deploying Jalapeño inside its own compute infrastructure by the end of the year, directly lowering the per-token power and latency costs of running ChatGPT-scale inference on open-weight models like GPT-OSS.
- Nvidia's GB200/GB300 systems lose their default status for frontier inference buyers: a credible second source now posts better efficiency numbers on exactly the models data centers actually serve.
Second-order effects
- Nvidia's counter-moves accelerate — the Groq LPU acquisition and Groq 3 LPX production ramp look like an admission that inference efficiency is the battleground, and sources' unconfirmed claim of system price increases of at least 15% starting in early 2027 would test whether customers defect rather than pay.
- Other large inference buyers face pressure to follow OpenAI's custom-silicon path or extract equivalent concessions from Nvidia, since Jalapeño's ~2 PB/s of HBM4 memory bandwidth (per analyst commentary on its 128-accelerator design) shows what purpose-built inference hardware can reach.
Third-order effects
- If Gen 2 — already described by OpenAI as deep in development — sustains this cadence, AI labs shift from renting merchant GPUs to owning multigenerational ASIC roadmaps, fragmenting the accelerator market Nvidia has dominated since the A100 era.
- Benchmark competition migrates from peak training flops to work-per-watt at serving latency, which favors players who control both the model and the chip — a structural edge that could harden around vertically integrated labs.
The trend: Frontier AI labs are verticalizing into their own inference silicon, making performance-per-watt on real serving workloads — not raw training throughput — the metric that decides accelerator-market share.
Related: Power-ready compute · Inference as a System Product · OpenAI · Nvidia · OpenAI and Broadcom unveil Jalapeño · SemiAnalysis details Jalapeño's win over Nvidia, AMD, and Google
Related Coverage
- OpenAI Jalapeño: Better Than Nvidia Blackwell SemiAnalysis · Bryan Shan
- OpenAI details Jalapeño AI chip, with 700W TDP DatacenterDynamics · Sebastian Moss
- OpenAI's Jalapeño chip is built for fast inference at scale, benchmarks show TechCrunch · Russell Brandom
- OpenAI says its Jalapeño chip offers spicy performance Axios · Ina Fried
- OpenAI's upcoming Jalapeño chip looks like it'll be an inference beast The Register
- OpenAI built a chip in nine months. Then it let AI rewrite the code. The New Stack · Amanda Caswell
- OpenAI Claims Its New Chips Can Outperform Nvidia Processors in Tests Bloomberg · Dina Bass
- Jalapeño's first results show industry-leading speed and efficiency in AI inference OpenAI
- OpenAI's Jalapeño spices up chip market as it outperforms Nvidia's Blackwell Seeking Alpha · Brandon Evans
- OpenAI says its Jalapeno AI chip delivers faster responses than rivals like Nvidia Digital Trends · Manisha Priyadarshini
- How OpenAI's Jalapeno chip just surprised the AI industry The Deep View · Jason Hiner
- The full stack behind abundant intelligence OpenAI · Sarah Friar
- OpenAI Unveils Jalapeño Chip with Industry-Leading AI Efficiency Blockchain.News · Rebeca Moen
- OpenAI's upcoming Jalapeño AI chip outperforms NVIDIA GB300 in inference tests Neowin · Pradeep Viswanathan
- OpenAI's Jalapeño chip claims major AI speed boost: What to know Livemint · Ravi Hari
- OpenAI: Jalapeño Chip Tops Efficiency Benchmarks Blockchain.News · Greg Brockman
- OpenAI's new AI chip outperforms Nvidia's GB300 in efficiency tests, company says Proactive · Angela Harmantas
- OpenAI Says New Jalapeno Chips Outperformed Nvidia in Testing Bloomberg
- OpenAI Releases First Benchmark Results for its Jalapeño AI Chip iPhone in Canada · Usman Qureshi
- OpenAI says Jalapeno beats Nvidia Blackwell on inference speed per watt RuntimeWire · Ryan Merket
- OpenAI's First-Gen Jalapeno ASIC Blows Competition Out Of The Park, Performs 1.5x to 1.9x More Work Per Kilowatt Than NVIDIA's Blackwell Chips, While Threatening The CUDA Moat Wccftech · Rohail Saleem
- OpenAI reports up to 3.6x lower Jalapeno latency in engineering tests RuntimeWire · Ryan Merket
- OpenAI's 700W Jalapeño ASIC outpaces 1,400W Nvidia flagship GPU — claims up to 1.9x throughput per kilowatt and 3.6x lower latency, co-developed with Broadcom Tom's Hardware · Luke James
- Wall Street Lunch: OpenAI's JalapeñO Tops Nvidia's Blackwell In Some Inference Tests Seeking Alpha
- OpenAI's first custom chip “Jalapeño” reportedly beats Nvidia's Blackwell and Rubin in inference benchmarks The Decoder · Matthias Bastian
- OpenAI Jalapeño: Better than Nvidia Blackwell Hacker News
- OpenAI Claims an AI Chip Breakthrough. Broadcom Collaboration Jalapeño Beats Nvidia's GB300 in Tests. International Business Times · Matias Civita
- First Benchmark Results on an OpenAI GPU TechTechPotato on YouTube
Discussion
-
@tobias_writes
Tobias Mann
on x
When it comes to inference, compute is key, but memory bandwidth is king. With 128 accelerators OpenAI's (& Broadcom) Jalapeño offers ~2 PB/s of HMB4 B/W — more than either Nvidia's Vera Rubin or AMD's Helios My 1,000+ word analysis only @TheRegister https://www.theregister.com/ …
-
@theahmadosman
Ahmad
on x
This is an important thing that happened today btw
-
@scaling01
@scaling01
on x
OpenAI says: “We plan to begin deploying Jalapeño within OpenAI's compute infrastructure by the end of the year. It is the first generation of a multigenerational roadmap: Gen 2 is deep in development, and Gen 3 is taking shape”
-
@dandr1s
Dan
on x
This is probably the most important Jalapeño graph. It measures the trade-off between: Interactivity: how fast each individual user gets tokens Utility: how much total AI output a data center gets per megawatt The ideal chip is in the top-right.
-
@dandr1s
Dan
on x
🚨 BREAKING: OpenAI's first custom AI chip is already beating Nvidia's latest systems on efficiency. Jalapeño was designed specifically to run models like ChatGPT and Codex, not train them. New benchmarks show: 1.5-1.9× more AI output per watt 1.7-3.6× lower latency than
-
@kimmonismus
@kimmonismus
on x
Holy: OpenAI says its first custom inference chip is already beating Nvidia GB200 and GB300 systems on speed and efficiency. Across GPT-OSS 120B, DeepSeek R1 670B, and Kimi K2.5 1T, Jalapeño delivered 1.5-1.9× more AI work per watt at peak throughput and 1.7-3.6× lower
-
@cpmou2022
Chengpeng
on x
some investors have been asking me lately about the wave of inference chip startups. My answer: we need real performance data before forming a view. So here's ours. Across three public models on InferenceX, Jalapeño delivered: • 1.5-1.9× more AI work per watt • 1.7-3.6× lower
-
@mweinbach
Max Weinbach
on x
Holy fuck OpenAI cooked https://openai.com/...
-
@scaling01
@scaling01
on x
guys it's beating VR200 NVL72
-
@rohanpaul_ai
Rohan Paul
on x
MASSIVE: OpenAI just claimed its new Jalapeño chips delivered 104.3x more throughput per kilowatt than NVIDIA GB300 at matched DeepSeek R1 decoding speed. The figure comes from GB300's previous-best time-between-token speed: at 169.41 tok/s/user, Jalapeño produced 12,258 mixed
-
@funofinvesting
Tevis
on x
OpenAI claims their Jalapeno chip (developed by $AVGO), which plans to deploy later this year, beat out $NVDA GB300s. My reaction: that's great, but kind of pointless unless you are comparing with the top of the line (it doesn't beat Vera Rubin).
-
@dylan522p
Dylan Patel
on x
OpenAI Jalapeno is spicy Usually first generation chips aren't competitive, but OpenAI is beating Nvidia Blackwell and even Rubin This is huge news! We got to go into OpenAI's lab to dissect their new chip and dive into the software, architecture, and performance Incredible work!
-
@jukan05
Jukan
on x
As far as I know, Jalapeño uses Samsung's HBM4 almost exclusively. It's superior to Rubin.
-
@zephyr_z9
@zephyr_z9
on x
OpenAI's chip design team cooked hard
-
@hakmgpt
@hakmgpt
on x
Wait wtf !!!! That's really huge results for new openAI chip ! It delivers higher performance per watt and lower latency than NVIDIA systems across multiple benchmarks, with compute, memory, and networking designed as one integrated system. Most remarkably, OpenAI used AI models
-
@realnickmugalli
Nicholas Mugalli
on x
OpenAI's custom inference chip, Jalapeño built alongside $AVGO just outpaced Nvidia's GB300 in internal tests on power efficiency and token speeds. This headline obviously focuses on the chip bypass (which hasn't been benchmarked against Nvidia's upcoming Vera Rubin), the real
-
@edludlow
Ed Ludlow
on x
OpenAI's published performance metrics for Jalapeno. Vs. Nvidia (Blackwell) in tests, OpenAI says Jalapeno led in two categories: the amount of AI work it could handle per unit of power and its speed at returning responses. OAI's Richard Ho is on Bloomberg Tech today
-
@ccatalini
Christian Catalini
on x
When ideas cannot be contained, control the atoms. @OpenAI needs a frontier chip to ensure value capture. https://www.a16z.news/...
-
@edzitron.com
Ed Zitron
on bluesky
Even in this example running its own chip on its own infrastructure, OpenAI refuses to actually document how much *cheaper* it would be to use its own chip, and doesn't even bother to benchmark on its own frontier models, defaulting to open source. Useless. [embedded post]
-
@zephyr_z9
@zephyr_z9
on x
Well, Nvidia has to sell 10M chips to lots of companies next year I'm pretty sure OpenAI will be super satisfied if 25%-35% of the compute deployed by them uses Jalapeno and its successors in 2028 and beyond They don't need/have to scale like Nvidia/Rubin
-
@gdb
Greg Brockman
on x
inference numbers published for jalapeno, team did an amazing job https://openai.com/...
-
@jukan05
Jukan
on x
One thing I can say is that OAI's Jalapeño can't scale to the same level as Rubin. If they source HBM4 exclusively from Samsung, there will be a limit to how far they can ramp. To scale beyond a certain point, they'll likely need to lower the HBM4 speed requirement so they can
-
@teortaxestex
@teortaxestex
on x
Jensen really has no purely technical moat anymore, huh at least on inference, frontier labs can cut him out
-
@firstadopter
Tae Kim
on x
OpenAI: “Today, we shared the first measured performance results from Jalapeño, OpenAI's first custom inference chip. On InferenceX, a public benchmark using GPT-OSS 120B, Jalapeño delivered more peak throughput per kilowatt and lower token latency than the commercial systems in
-
r/NVDA_Stock
r
on reddit
OpenAI' Jalapeño: Better Than Nvidia Blackwell
-
@suchenzang
Susan Zhang
on x
hmmm (this is still a super cool release either way, congrats to the ex TPU team :))
-
@cgtwts
@cgtwts
on x
OpenAI's first custom chip preview is already beating Nvidia's GB300 🤯
-
@openai
@openai
on x
Since announcing Jalapeño, our first custom inference chip, we've been testing it and the system around it. The results show a major advance: more intelligence from every watt and faster responses, delivering both higher throughput and lower latency in one architecture without
-
@cafkafk
@cafkafk
on x
No way they called that shit Jalapeño and it seems like its going pretty hard ngl. https://openai.com/...
-
@anujsaharan
Anuj Saharan
on x
good morning, the gang has made some spicy chips for you https://openai.com/...
-
@scaling01
@scaling01
on x
it's more interesting that OpenAI had Astra for ~2 months
-
@itsclivetime
Clive Chan
on x
Jalapeno beats VR200 in A0 silicon with a vibe-ported, non-specdec implementation of DeepSeek 🌶️ 🌶️ And with much faster program execution than VR200, and with a B0 update landing imminently Congrats to the Jalapeno team!!!
-
@youjiacheng
You Jiacheng
on x
No MTP, No PD disaggregation, Pure TP, still beats NVIDIA's Vera Rubin NVL72 on a third-party model, with A0 stepping. And B0 is 25% better. NVIDIA GPUs become HBM wrappers.
-
@liamfedus
Liam Fedus
on x
Data centers are power-limited and OpenAI's Jalapeño sets new frontiers for tokens per second per watt. Impressive and fast work from OpenAI's ASIC program.
-
@eladgil
Elad Gil
on x
Impressive speed to tape out by OpenAI https://newsletter.semianalysis.com/ ...
-
@jordannanos
Jordan Nanos
on x
OpenAI is designing for perf/W because they are currently limited by datacenter power envelopes (not budget or floor space) and have produced the leading accelerator in the industry by that metric It took 9 months to go from initial RTL to tapeout, and 6-8 more months for it to
-
@eliebakouch
Elie
on x
> Using Codex with GPT-Astra, the team brought three open-weight models that were not part of Jalapeño's original production plan to high performance within two months from openai blog, two months is quite a lot here no? i'm a bit confuse by this
-
@eliebakouch
Elie
on x
1. openai doesn't use MLA for attention mechanism 2. they implemented MLA kernel with codex without any human supervision
-
@noslawextratost
Cam Quilici
on x
One of the most interesting parts of this is how the Jalapeno team (allegedly, and I think this is reasonable) used AI to accelerate the bringup of the chip on frontier(ish) models They can literally just profile the runs and hand it off to Codex for infinite iteration This
-
@thelokasiffers
Giulio
on x
maybe I'm retarded but isn't it expected for an ASIC to outperform a general-purpose GPU on the workload the ASIC was specifically designed for?
-
@cdleary
@cdleary
on x
it's a good chip ser agree w/ the article: the team is cracked and delivered like you wouldn't believe (gonna be a lot of cope) the spice must flow 🌶️
-
@semianalysis_
@semianalysis_
on x
OpenAI Jalapeño: Better Than Nvidia Blackwell OpenAI's self-designed ASIC compared with Rubin, Jalapeño's TCO, throughput per MW, and spicy deets https://newsletter.semianalysis.com/ ...
-
Sarah Friar
Sarah Friar
on linkedin
Chip drop. 🎤 🫳 — This is what investing ahead looks like in practice. — Today, OpenAI shared the first measured results …
-
Richard Ho
Richard Ho
on linkedin
🌶️ Jalapeño's first detailed performance results show industry-leading speed and efficiency for AI inference. …
-
r/singularity
r
on reddit
OpenAI blog post on their new custom inference chip
-
Madeleine Laitz, PhD
Madeleine Laitz, PhD
on linkedin
Really excited to see Jalapeño out in the world from OpenAI. The most interesting part to me is the strategy …
-
@sama
Sam Altman
on x
we made a chip and it is fast
-
Sachin Katti
Sachin Katti
on linkedin
OpenAI's compute strategy requires both breadth and leverage: deep partnerships with the industry's leading accelerator, cloud, systems …
-
r/OpenAI
r
on reddit
OpenAI Jalapeño: Better Than Nvidia Blackwell
-
r/AMD_Stock
r
on reddit
OpenAI' Jalapeño: Better Than Nvidia Blackwell
-
@tszzl
Roon
on x
ultrafast inference reveals new threats. worth thinking about how quickly misaligned frontier class models running 50x faster could infiltrate systems and so on— far too quickly for human responders to stay abreast. you need autonomous detection and shutdown, not just monitoring
-
@bubbleboi
Bubble Boi
on x
My favorite technical detail about jalapeño is that although OpenAI designed it with Broadcom. Broadcom is lending them money to manufacture it since you can't get a loan for the NRE for taping out a chip. Makes sense then why Broadcom went so hard and made sure it was good
-
@thsottiaux
Tibo
on x
Excited about our Jalapeno results today. An incredible achievement from the team, taking a new chip from concept all the way to very impressive performance on real workloads in the lab. Tomorrow's fast will feel like today's ultrafast. As I've mentioned before, we're pushing to
-
@openai
@openai
on x
Since announcing Jalapeño, our first custom inference chip, we've been testing it and the system around it. The results show a major advance: more intelligence from every watt and faster responses, delivering both higher throughput and lower latency in one architecture without
-
@maxkan
Max Kan
on x
Jalepeno is the best example we have so far of extremely high roi token spend by a large enterprise to create an entirely new business line. No surprise that it came first from OpenAI/Anthropic.
-
@patrickmoorhead
Patrick Moorhead
on x
I will be amazed if 🌶️ is as good in production at scale as it is on the simple benchmarks shown today. I've spent over 30 years in chips and I've never seen a decent V1 do well at scale on production workloads. Maybe there was a V.5. Maybe it won't provide the performance in
-
@edzitron
Ed Zitron
on x
It's hilarious that even with their own chip running in their own infrastructure OpenAI refuses to make any statements about how its chip would run its own closed-source models, just more open source comparisons. ChatGPT isn't running GPT-OSS buddy