/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

← → days · ↑ ↓ browse · Enter similar · o open

OpenAI says its Jalapeño chip delivered 1.5x-1.9x more AI work per watt and 1.7x-3.6x lower latency than Nvidia chips across GPT-OSS, DeepSeek R1, Kimi K2.5 1T

Jalapeño outperformed Nvidia's superchips on an AI inference benchmark test.

The Verge Emma Roth

Context & Ripple Effects

When OpenAI and Broadcom first unveiled Jalapeño in June, the claim was process speed — an LLM-optimized inference chip taken from design to manufacturing tape-out in about nine months. Today's benchmarks are the payoff: across GPT-OSS, DeepSeek R1, and Kimi K2.5 1T, OpenAI reports 1.5x-1.9x more work per watt and 1.7x-3.6x lower latency than Nvidia's superchips, and a SemiAnalysis deep-dive published alongside adds that the ASIC also beat AMD and Google parts on several top open-weight models.

The timing sharpens the rivalry: Nvidia spent the summer courting OpenAI as a flagship customer — Jensen Huang named it among the first big users of the new Vera CPUs in June — while separately confirming a $20 billion bet on Groq's LPU technology and putting its Groq 3 LPX inference accelerator into full production. OpenAI is answering not with a purchase but with silicon of its own.

First-order effects

  • OpenAI says it will begin deploying Jalapeño inside its own compute infrastructure by the end of the year, directly lowering the per-token power and latency costs of running ChatGPT-scale inference on open-weight models like GPT-OSS.
  • Nvidia's GB200/GB300 systems lose their default status for frontier inference buyers: a credible second source now posts better efficiency numbers on exactly the models data centers actually serve.

Second-order effects

  • Nvidia's counter-moves accelerate — the Groq LPU acquisition and Groq 3 LPX production ramp look like an admission that inference efficiency is the battleground, and sources' unconfirmed claim of system price increases of at least 15% starting in early 2027 would test whether customers defect rather than pay.
  • Other large inference buyers face pressure to follow OpenAI's custom-silicon path or extract equivalent concessions from Nvidia, since Jalapeño's ~2 PB/s of HBM4 memory bandwidth (per analyst commentary on its 128-accelerator design) shows what purpose-built inference hardware can reach.

Third-order effects

  • If Gen 2 — already described by OpenAI as deep in development — sustains this cadence, AI labs shift from renting merchant GPUs to owning multigenerational ASIC roadmaps, fragmenting the accelerator market Nvidia has dominated since the A100 era.
  • Benchmark competition migrates from peak training flops to work-per-watt at serving latency, which favors players who control both the model and the chip — a structural edge that could harden around vertically integrated labs.

The trend: Frontier AI labs are verticalizing into their own inference silicon, making performance-per-watt on real serving workloads — not raw training throughput — the metric that decides accelerator-market share.

Discussion

  • @tobias_writes Tobias Mann on x
    When it comes to inference, compute is key, but memory bandwidth is king. With 128 accelerators OpenAI's (& Broadcom) Jalapeño offers ~2 PB/s of HMB4 B/W — more than either Nvidia's Vera Rubin or AMD's Helios My 1,000+ word analysis only @TheRegister https://www.theregister.com/ …
  • @theahmadosman Ahmad on x
    This is an important thing that happened today btw
  • @scaling01 @scaling01 on x
    OpenAI says: “We plan to begin deploying Jalapeño within OpenAI's compute infrastructure by the end of the year. It is the first generation of a multigenerational roadmap: Gen 2 is deep in development, and Gen 3 is taking shape”
  • @dandr1s Dan on x
    This is probably the most important Jalapeño graph. It measures the trade-off between: Interactivity: how fast each individual user gets tokens Utility: how much total AI output a data center gets per megawatt The ideal chip is in the top-right.
  • @dandr1s Dan on x
    🚨 BREAKING: OpenAI's first custom AI chip is already beating Nvidia's latest systems on efficiency. Jalapeño was designed specifically to run models like ChatGPT and Codex, not train them. New benchmarks show: 1.5-1.9× more AI output per watt 1.7-3.6× lower latency than
  • @kimmonismus @kimmonismus on x
    Holy: OpenAI says its first custom inference chip is already beating Nvidia GB200 and GB300 systems on speed and efficiency. Across GPT-OSS 120B, DeepSeek R1 670B, and Kimi K2.5 1T, Jalapeño delivered 1.5-1.9× more AI work per watt at peak throughput and 1.7-3.6× lower
  • @cpmou2022 Chengpeng on x
    some investors have been asking me lately about the wave of inference chip startups. My answer: we need real performance data before forming a view. So here's ours. Across three public models on InferenceX, Jalapeño delivered: • 1.5-1.9× more AI work per watt • 1.7-3.6× lower
  • @mweinbach Max Weinbach on x
    Holy fuck OpenAI cooked https://openai.com/...
  • @scaling01 @scaling01 on x
    guys it's beating VR200 NVL72
  • @rohanpaul_ai Rohan Paul on x
    MASSIVE: OpenAI just claimed its new Jalapeño chips delivered 104.3x more throughput per kilowatt than NVIDIA GB300 at matched DeepSeek R1 decoding speed. The figure comes from GB300's previous-best time-between-token speed: at 169.41 tok/s/user, Jalapeño produced 12,258 mixed
  • @funofinvesting Tevis on x
    OpenAI claims their Jalapeno chip (developed by $AVGO), which plans to deploy later this year, beat out $NVDA GB300s. My reaction: that's great, but kind of pointless unless you are comparing with the top of the line (it doesn't beat Vera Rubin).
  • @dylan522p Dylan Patel on x
    OpenAI Jalapeno is spicy Usually first generation chips aren't competitive, but OpenAI is beating Nvidia Blackwell and even Rubin This is huge news! We got to go into OpenAI's lab to dissect their new chip and dive into the software, architecture, and performance Incredible work!
  • @jukan05 Jukan on x
    As far as I know, Jalapeño uses Samsung's HBM4 almost exclusively. It's superior to Rubin.
  • @zephyr_z9 @zephyr_z9 on x
    OpenAI's chip design team cooked hard
  • @hakmgpt @hakmgpt on x
    Wait wtf !!!! That's really huge results for new openAI chip ! It delivers higher performance per watt and lower latency than NVIDIA systems across multiple benchmarks, with compute, memory, and networking designed as one integrated system. Most remarkably, OpenAI used AI models
  • @realnickmugalli Nicholas Mugalli on x
    OpenAI's custom inference chip, Jalapeño built alongside $AVGO just outpaced Nvidia's GB300 in internal tests on power efficiency and token speeds. This headline obviously focuses on the chip bypass (which hasn't been benchmarked against Nvidia's upcoming Vera Rubin), the real
  • @edludlow Ed Ludlow on x
    OpenAI's published performance metrics for Jalapeno. Vs. Nvidia (Blackwell) in tests, OpenAI says Jalapeno led in two categories: the amount of AI work it could handle per unit of power and its speed at returning responses. OAI's Richard Ho is on Bloomberg Tech today
  • @ccatalini Christian Catalini on x
    When ideas cannot be contained, control the atoms. @OpenAI needs a frontier chip to ensure value capture. https://www.a16z.news/...
  • @edzitron.com Ed Zitron on bluesky
    Even in this example running its own chip on its own infrastructure, OpenAI refuses to actually document how much *cheaper* it would be to use its own chip, and doesn't even bother to benchmark on its own frontier models, defaulting to open source.  Useless.  [embedded post]
  • @zephyr_z9 @zephyr_z9 on x
    Well, Nvidia has to sell 10M chips to lots of companies next year I'm pretty sure OpenAI will be super satisfied if 25%-35% of the compute deployed by them uses Jalapeno and its successors in 2028 and beyond They don't need/have to scale like Nvidia/Rubin
  • @gdb Greg Brockman on x
    inference numbers published for jalapeno, team did an amazing job https://openai.com/...
  • @jukan05 Jukan on x
    One thing I can say is that OAI's Jalapeño can't scale to the same level as Rubin. If they source HBM4 exclusively from Samsung, there will be a limit to how far they can ramp. To scale beyond a certain point, they'll likely need to lower the HBM4 speed requirement so they can
  • @teortaxestex @teortaxestex on x
    Jensen really has no purely technical moat anymore, huh at least on inference, frontier labs can cut him out
  • @firstadopter Tae Kim on x
    OpenAI: “Today, we shared the first measured performance results from Jalapeño, OpenAI's first custom inference chip. On InferenceX, a public benchmark using GPT-OSS 120B, Jalapeño delivered more peak throughput per kilowatt and lower token latency than the commercial systems in
  • r/NVDA_Stock r on reddit
    OpenAI' Jalapeño: Better Than Nvidia Blackwell
  • @suchenzang Susan Zhang on x
    hmmm (this is still a super cool release either way, congrats to the ex TPU team :))
  • @cgtwts @cgtwts on x
    OpenAI's first custom chip preview is already beating Nvidia's GB300 🤯
  • @openai @openai on x
    Since announcing Jalapeño, our first custom inference chip, we've been testing it and the system around it. The results show a major advance: more intelligence from every watt and faster responses, delivering both higher throughput and lower latency in one architecture without
  • @cafkafk @cafkafk on x
    No way they called that shit Jalapeño and it seems like its going pretty hard ngl. https://openai.com/...
  • @anujsaharan Anuj Saharan on x
    good morning, the gang has made some spicy chips for you https://openai.com/...
  • @scaling01 @scaling01 on x
    it's more interesting that OpenAI had Astra for ~2 months
  • @itsclivetime Clive Chan on x
    Jalapeno beats VR200 in A0 silicon with a vibe-ported, non-specdec implementation of DeepSeek 🌶️ 🌶️ And with much faster program execution than VR200, and with a B0 update landing imminently Congrats to the Jalapeno team!!!
  • @youjiacheng You Jiacheng on x
    No MTP, No PD disaggregation, Pure TP, still beats NVIDIA's Vera Rubin NVL72 on a third-party model, with A0 stepping. And B0 is 25% better. NVIDIA GPUs become HBM wrappers.
  • @liamfedus Liam Fedus on x
    Data centers are power-limited and OpenAI's Jalapeño sets new frontiers for tokens per second per watt. Impressive and fast work from OpenAI's ASIC program.
  • @eladgil Elad Gil on x
    Impressive speed to tape out by OpenAI https://newsletter.semianalysis.com/ ...
  • @jordannanos Jordan Nanos on x
    OpenAI is designing for perf/W because they are currently limited by datacenter power envelopes (not budget or floor space) and have produced the leading accelerator in the industry by that metric It took 9 months to go from initial RTL to tapeout, and 6-8 more months for it to
  • @eliebakouch Elie on x
    > Using Codex with GPT-Astra, the team brought three open-weight models that were not part of Jalapeño's original production plan to high performance within two months from openai blog, two months is quite a lot here no? i'm a bit confuse by this
  • @eliebakouch Elie on x
    1. openai doesn't use MLA for attention mechanism 2. they implemented MLA kernel with codex without any human supervision
  • @noslawextratost Cam Quilici on x
    One of the most interesting parts of this is how the Jalapeno team (allegedly, and I think this is reasonable) used AI to accelerate the bringup of the chip on frontier(ish) models They can literally just profile the runs and hand it off to Codex for infinite iteration This
  • @thelokasiffers Giulio on x
    maybe I'm retarded but isn't it expected for an ASIC to outperform a general-purpose GPU on the workload the ASIC was specifically designed for?
  • @cdleary @cdleary on x
    it's a good chip ser agree w/ the article: the team is cracked and delivered like you wouldn't believe (gonna be a lot of cope) the spice must flow 🌶️
  • @semianalysis_ @semianalysis_ on x
    OpenAI Jalapeño: Better Than Nvidia Blackwell OpenAI's self-designed ASIC compared with Rubin, Jalapeño's TCO, throughput per MW, and spicy deets https://newsletter.semianalysis.com/ ...
  • Sarah Friar Sarah Friar on linkedin
    Chip drop.  🎤 🫳  —  This is what investing ahead looks like in practice.  —  Today, OpenAI shared the first measured results …
  • Richard Ho Richard Ho on linkedin
    🌶️ Jalapeño's first detailed performance results show industry-leading speed and efficiency for AI inference. …
  • r/singularity r on reddit
    OpenAI blog post on their new custom inference chip
  • Madeleine Laitz, PhD Madeleine Laitz, PhD on linkedin
    Really excited to see Jalapeño out in the world from OpenAI.  The most interesting part to me is the strategy …
  • @sama Sam Altman on x
    we made a chip and it is fast
  • Sachin Katti Sachin Katti on linkedin
    OpenAI's compute strategy requires both breadth and leverage: deep partnerships with the industry's leading accelerator, cloud, systems …
  • r/OpenAI r on reddit
    OpenAI Jalapeño: Better Than Nvidia Blackwell
  • r/AMD_Stock r on reddit
    OpenAI' Jalapeño: Better Than Nvidia Blackwell
  • @tszzl Roon on x
    ultrafast inference reveals new threats. worth thinking about how quickly misaligned frontier class models running 50x faster could infiltrate systems and so on— far too quickly for human responders to stay abreast. you need autonomous detection and shutdown, not just monitoring
  • @bubbleboi Bubble Boi on x
    My favorite technical detail about jalapeño is that although OpenAI designed it with Broadcom. Broadcom is lending them money to manufacture it since you can't get a loan for the NRE for taping out a chip. Makes sense then why Broadcom went so hard and made sure it was good
  • @thsottiaux Tibo on x
    Excited about our Jalapeno results today. An incredible achievement from the team, taking a new chip from concept all the way to very impressive performance on real workloads in the lab. Tomorrow's fast will feel like today's ultrafast. As I've mentioned before, we're pushing to
  • @openai @openai on x
    Since announcing Jalapeño, our first custom inference chip, we've been testing it and the system around it. The results show a major advance: more intelligence from every watt and faster responses, delivering both higher throughput and lower latency in one architecture without
  • @maxkan Max Kan on x
    Jalepeno is the best example we have so far of extremely high roi token spend by a large enterprise to create an entirely new business line. No surprise that it came first from OpenAI/Anthropic.
  • @patrickmoorhead Patrick Moorhead on x
    I will be amazed if 🌶️ is as good in production at scale as it is on the simple benchmarks shown today. I've spent over 30 years in chips and I've never seen a decent V1 do well at scale on production workloads. Maybe there was a V.5. Maybe it won't provide the performance in
  • @edzitron Ed Zitron on x
    It's hilarious that even with their own chip running in their own infrastructure OpenAI refuses to make any statements about how its chip would run its own closed-source models, just more open source comparisons. ChatGPT isn't running GPT-OSS buddy