/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

A detailed look at Jalapeño, OpenAI's ASIC developed with Broadcom in 16 months, which beat Nvidia, AMD, and Google chips on multiple top open-weight models

OpenAI's self-designed ASIC compared with Rubin, Jalapeño's TCO, throughput per MW, and spicy deets

SemiAnalysis

Context & Ripple Effects

The arc here runs from finance to silicon to benchmarks: Bloomberg reported in October 2025 that OpenAI expected to spend 20% to 30% less on co-developed Broadcom chips than on backlogged Nvidia GPUs, then OpenAI and Broadcom unveiled Jalapeño in June 2026 with a nine-month design-to-tape-out cycle. The SemiAnalysis deep-dive now puts the program at 16 months end-to-end and publishes comparative numbers across GPT-OSS, DeepSeek R1, and Kimi K2.5 1T.

The story travelled unusually wide on pickup day — Bloomberg, Tom's Hardware, ServeTheHome, Seeking Alpha and others all ran it alongside OpenAI's own 'full stack' post and a Hot Chips 2026 presentation. Reception split along predictable lines: Elad Gil flagged the tape-out speed, while commenters like @teortaxestex argued Jensen Huang has 'no purely technical moat' left on inference, and skeptics noted an ASIC beating general-purpose GPUs on its own target workload is close to expected.

First-order effects

  • OpenAI's own inference fleet becomes the reference deployment: at a claimed 700W versus a 1,400W Nvidia flagship, the chip's throughput-per-MW economics apply first to OpenAI's serving costs for GPT-OSS and partner models, validating the cost-reduction thesis it briefed Bloomberg on last year.
  • Nvidia, AMD, and Google now have a published, third-party-analyzed benchmark series where their top parts lose on work-per-watt and latency for LLM inference — a marketing and procurement battleground they must answer in kind rather than ignore.

Second-order effects

  • Nvidia's counter-move is already visible in the corpus: its Groq 3 LPX AI inference accelerator entering full production as of late August 2026 signals a dedicated-inference product line responding to exactly this ASIC threat, rather than ceding the segment to general-purpose GPUs.
  • Broadcom emerges as the kingmaker supplier: if the Jalapeño playbook (lab-owned architecture, Broadcom manufacturing muscle, models accelerating the design loop) is repeatable in 16 months, every capital-rich frontier lab faces pressure to commission its own part, straining Broadcom's design-services capacity and tightening supply of advanced packaging.

Third-order effects

  • If the pattern holds, inference compute splits into a two-tier market: vertically integrated labs running proprietary ASICs at superior cost-per-token, and everyone else renting merchant GPUs — mirroring the hyperscaler TPU trajectory and eroding Nvidia's inference pricing power over the largest customers.
  • Vendor-self-published benchmarks become the norm for custom silicon, pushing buyers and analysts toward independent verification as the deciding factor — the credibility of these 1.5x-3.6x claims rests on whether neutral parties can reproduce them at scale.

The trend: Frontier AI labs are following hyperscalers into workload-specific silicon, using their own models to compress design cycles and attacking Nvidia where it is most exposed — high-volume inference.

Discussion

  • @openai @openai on x
    Since announcing Jalapeño, our first custom inference chip, we've been testing it and the system around it. The results show a major advance: more intelligence from every watt and faster responses, delivering both higher throughput and lower latency in one architecture without
  • @eliebakouch Elie on x
    > Using Codex with GPT-Astra, the team brought three open-weight models that were not part of Jalapeño's original production plan to high performance within two months from openai blog, two months is quite a lot here no? i'm a bit confuse by this
  • @teortaxestex @teortaxestex on x
    Jensen really has no purely technical moat anymore, huh at least on inference, frontier labs can cut him out
  • @gdb Greg Brockman on x
    inference numbers published for jalapeno, team did an amazing job https://openai.com/...
  • @eladgil Elad Gil on x
    Impressive speed to tape out by OpenAI https://newsletter.semianalysis.com/ ...
  • @thelokasiffers Giulio on x
    maybe I'm retarded but isn't it expected for an ASIC to outperform a general-purpose GPU on the workload the ASIC was specifically designed for?
  • @maxkan Max Kan on x
    Jalepeno is the best example we have so far of extremely high roi token spend by a large enterprise to create an entirely new business line. No surprise that it came first from OpenAI/Anthropic.
  • @scaling01 @scaling01 on x
    it's more interesting that OpenAI had Astra for ~2 months
  • @youjiacheng You Jiacheng on x
    No MTP, No PD disaggregation, Pure TP, still beats NVIDIA's Vera Rubin NVL72 on a third-party model, with A0 stepping. And B0 is 25% better. NVIDIA GPUs become HBM wrappers.
  • @zephyr_z9 @zephyr_z9 on x
    Well, Nvidia has to sell 10M chips to lots of companies next year I'm pretty sure OpenAI will be super satisfied if 25%-35% of the compute deployed by them uses Jalapeno and its successors in 2028 and beyond They don't need/have to scale like Nvidia/Rubin
  • @cdleary @cdleary on x
    it's a good chip ser agree w/ the article: the team is cracked and delivered like you wouldn't believe (gonna be a lot of cope) the spice must flow 🌶️
  • @jordannanos Jordan Nanos on x
    OpenAI is designing for perf/W because they are currently limited by datacenter power envelopes (not budget or floor space) and have produced the leading accelerator in the industry by that metric It took 9 months to go from initial RTL to tapeout, and 6-8 more months for it to
  • @liamfedus Liam Fedus on x
    Data centers are power-limited and OpenAI's Jalapeño sets new frontiers for tokens per second per watt. Impressive and fast work from OpenAI's ASIC program.
  • @itsclivetime Clive Chan on x
    Jalapeno beats VR200 in A0 silicon with a vibe-ported, non-specdec implementation of DeepSeek 🌶️ 🌶️ And with much faster program execution than VR200, and with a B0 update landing imminently Congrats to the Jalapeno team!!!
  • @firstadopter Tae Kim on x
    OpenAI: “Today, we shared the first measured performance results from Jalapeño, OpenAI's first custom inference chip. On InferenceX, a public benchmark using GPT-OSS 120B, Jalapeño delivered more peak throughput per kilowatt and lower token latency than the commercial systems in
  • @semianalysis_ @semianalysis_ on x
    OpenAI Jalapeño: Better Than Nvidia Blackwell OpenAI's self-designed ASIC compared with Rubin, Jalapeño's TCO, throughput per MW, and spicy deets https://newsletter.semianalysis.com/ ...
  • @eliebakouch Elie on x
    1. openai doesn't use MLA for attention mechanism 2. they implemented MLA kernel with codex without any human supervision
  • @noslawextratost Cam Quilici on x
    One of the most interesting parts of this is how the Jalapeno team (allegedly, and I think this is reasonable) used AI to accelerate the bringup of the chip on frontier(ish) models They can literally just profile the runs and hand it off to Codex for infinite iteration This
  • Madeleine Laitz, PhD Madeleine Laitz, PhD on linkedin
    Really excited to see Jalapeño out in the world from OpenAI.  The most interesting part to me is the strategy …
  • r/OpenAI r on reddit
    OpenAI Jalapeño: Better Than Nvidia Blackwell
  • r/AMD_Stock r on reddit
    OpenAI' Jalapeño: Better Than Nvidia Blackwell
  • r/NVDA_Stock r on reddit
    OpenAI' Jalapeño: Better Than Nvidia Blackwell
  • @jukan05 Jukan on x
    What Jalapeño made me realize is just how much lower the barrier to chip design is going to become.  Ultimately, manufacturing will be where the real moat remains—at least until we have robots capable enough to change that.  Design and engineering outside of manufacturing will gr…
  • @andrewcurran_ Andrew Curran on x
    GPT-Astra helped develop the Jalapeño chip. 'Supporting each new model family still requires new kernels and model-specific optimizations. Using Codex with GPT-Astra, the team brought three open-weight models that were not part of Jalapeño's original production plan to high
  • @petergostev Peter Gostev on x
    I don't particularly believe we are in an automated research era, but I totally believe that engineering powered by AI will be eating the world
  • @patricktoulme Patrick C Toulme on x
    OpenAI Jalapeño is truly the first AI silicon developed by GPT-Astra and other OpenAI internal models. I believe they heavily used reinforcement learning on internal models like GPT-Astra to achieve these SOTA results. Here is what I think they did: 1. RL an internal GPT to
  • @sama Sam Altman on x
    we made a chip and it is fast
  • @dylan522p Dylan Patel on x
    OpenAI Jalapeno is spicy Usually first generation chips aren't competitive, but OpenAI is beating Nvidia Blackwell and even Rubin This is huge news! We got to go into OpenAI's lab to dissect their new chip and dive into the software, architecture, and performance Incredible work!
  • @tszzl Roon on x
    ultrafast inference reveals new threats. worth thinking about how quickly misaligned frontier class models running 50x faster could infiltrate systems and so on— far too quickly for human responders to stay abreast. you need autonomous detection and shutdown, not just monitoring
  • @patrickmoorhead Patrick Moorhead on x
    I will be amazed if 🌶️ is as good in production at scale as it is on the simple benchmarks shown today.  I've spent over 30 years in chips and I've never seen a decent V1 do well at scale on production workloads.  Maybe there was a V.5.  Maybe it won't provide the performance in …
  • @cerebras @cerebras on x
    Congratulations to our partners at @OpenAI today.  Taking a new chip to real lab results at great performance is a huge achievement, and improving speed across the board pushes our whole industry forward.  We look forward to having Jalapeno and the next-gen Cerebras solution toge…
  • @milksandmatcha Sarah Chieng on x
    I don't think people realize what Jalapeno + Cerebras is going to unlock... it's another abacus moment
  • @bubbleboi Bubble Boi on x
    My favorite technical detail about jalapeño is that although OpenAI designed it with Broadcom.  Broadcom is lending them money to manufacture it since you can't get a loan for the NRE for taping out a chip.  Makes sense then why Broadcom went so hard and made sure it was good cau…
  • @thsottiaux Tibo on x
    Excited about our Jalapeno results today.  An incredible achievement from the team, taking a new chip from concept all the way to very impressive performance on real workloads in the lab.  Tomorrow's fast will feel like today's ultrafast.  As I've mentioned before, we're pushing …
  • @bubbleboi Bubble Boi on x
    Big reason we aren't seeing Broadcom move higher on the jalaleño news imo. 1. Already priced in. OpenAI guides to ~30 GW by 2030 and assuming Jalapeño is 30-40%, that's 9-12 GW by 2030 which already matches the 10 GW Broadcom deal signed in October 2025 2. Content per GW is
  • @suchenzang Susan Zhang on x
    so i'm guessing the markets are seeing some diminishing returns wrt scale here? either way, congrats to the ex TPU team for shipping jalapeno! was going to QT the official announcement for engagement bait but alas, ran out of patience.
  • @suchenzang Susan Zhang on x
    hmmm (this is still a super cool release either way, congrats to the ex TPU team :))
  • @cgtwts @cgtwts on x
    OpenAI's first custom chip preview is already beating Nvidia's GB300 🤯
  • @openai @openai on x
    Since announcing Jalapeño, our first custom inference chip, we've been testing it and the system around it. The results show a major advance: more intelligence from every watt and faster responses, delivering both higher throughput and lower latency in one architecture without
  • @jukan05 Jukan on x
    One thing I can say is that OAI's Jalapeño can't scale to the same level as Rubin.  If they source HBM4 exclusively from Samsung, there will be a limit to how far they can ramp.  To scale beyond a certain point, they'll likely need to lower the HBM4 speed requirement so they can …
  • @gavinsbaker Gavin Baker on x
    Impressive that Jalapeño outperforms the comparable TPU and is in the mix with Rubin.  Credit where credit is due - first good ASIC outside of TPU/Trainium.  However, will likely significantly underperform a disaggregated GPU/Trainum plus SRAM accelerator setup.  Especially with …
  • @cafkafk @cafkafk on x
    No way they called that shit Jalapeño and it seems like its going pretty hard ngl. https://openai.com/...
  • @ccatalini Christian Catalini on x
    When ideas cannot be contained, control the atoms. @OpenAI needs a frontier chip to ensure value capture. https://www.a16z.news/...
  • @anujsaharan Anuj Saharan on x
    good morning, the gang has made some spicy chips for you https://openai.com/...
  • @tobias_writes Tobias Mann on x
    When it comes to inference, compute is key, but memory bandwidth is king. With 128 accelerators OpenAI's (& Broadcom) Jalapeño offers ~2 PB/s of HMB4 B/W — more than either Nvidia's Vera Rubin or AMD's Helios My 1,000+ word analysis only @TheRegister https://www.theregister.com/ …
  • @theahmadosman Ahmad on x
    This is an important thing that happened today btw
  • @scaling01 @scaling01 on x
    OpenAI says: “We plan to begin deploying Jalapeño within OpenAI's compute infrastructure by the end of the year. It is the first generation of a multigenerational roadmap: Gen 2 is deep in development, and Gen 3 is taking shape”
  • @edzitron Ed Zitron on x
    It's hilarious that even with their own chip running in their own infrastructure OpenAI refuses to make any statements about how its chip would run its own closed-source models, just more open source comparisons. ChatGPT isn't running GPT-OSS buddy
  • @dandr1s Dan on x
    This is probably the most important Jalapeño graph. It measures the trade-off between: Interactivity: how fast each individual user gets tokens Utility: how much total AI output a data center gets per megawatt The ideal chip is in the top-right.
  • @dandr1s Dan on x
    🚨 BREAKING: OpenAI's first custom AI chip is already beating Nvidia's latest systems on efficiency. Jalapeño was designed specifically to run models like ChatGPT and Codex, not train them. New benchmarks show: 1.5-1.9× more AI output per watt 1.7-3.6× lower latency than
  • @kimmonismus @kimmonismus on x
    Holy: OpenAI says its first custom inference chip is already beating Nvidia GB200 and GB300 systems on speed and efficiency. Across GPT-OSS 120B, DeepSeek R1 670B, and Kimi K2.5 1T, Jalapeño delivered 1.5-1.9× more AI work per watt at peak throughput and 1.7-3.6× lower
  • @cpmou2022 Chengpeng on x
    some investors have been asking me lately about the wave of inference chip startups. My answer: we need real performance data before forming a view. So here's ours. Across three public models on InferenceX, Jalapeño delivered: • 1.5-1.9× more AI work per watt • 1.7-3.6× lower
  • @mweinbach Max Weinbach on x
    Holy fuck OpenAI cooked https://openai.com/...
  • @scaling01 @scaling01 on x
    guys it's beating VR200 NVL72
  • @rohanpaul_ai Rohan Paul on x
    MASSIVE: OpenAI just claimed its new Jalapeño chips delivered 104.3x more throughput per kilowatt than NVIDIA GB300 at matched DeepSeek R1 decoding speed. The figure comes from GB300's previous-best time-between-token speed: at 169.41 tok/s/user, Jalapeño produced 12,258 mixed
  • @funofinvesting Tevis on x
    OpenAI claims their Jalapeno chip (developed by $AVGO), which plans to deploy later this year, beat out $NVDA GB300s. My reaction: that's great, but kind of pointless unless you are comparing with the top of the line (it doesn't beat Vera Rubin).
  • @jukan05 Jukan on x
    As far as I know, Jalapeño uses Samsung's HBM4 almost exclusively. It's superior to Rubin.
  • @zephyr_z9 @zephyr_z9 on x
    OpenAI's chip design team cooked hard
  • @hakmgpt @hakmgpt on x
    Wait wtf !!!! That's really huge results for new openAI chip ! It delivers higher performance per watt and lower latency than NVIDIA systems across multiple benchmarks, with compute, memory, and networking designed as one integrated system. Most remarkably, OpenAI used AI models
  • @realnickmugalli Nicholas Mugalli on x
    OpenAI's custom inference chip, Jalapeño built alongside $AVGO just outpaced Nvidia's GB300 in internal tests on power efficiency and token speeds. This headline obviously focuses on the chip bypass (which hasn't been benchmarked against Nvidia's upcoming Vera Rubin), the real
  • @edludlow Ed Ludlow on x
    OpenAI's published performance metrics for Jalapeno. Vs. Nvidia (Blackwell) in tests, OpenAI says Jalapeno led in two categories: the amount of AI work it could handle per unit of power and its speed at returning responses. OAI's Richard Ho is on Bloomberg Tech today
  • Sarah Friar Sarah Friar on linkedin
    Chip drop.  🎤 🫳  —  This is what investing ahead looks like in practice.  —  Today, OpenAI shared the first measured results …
  • Sachin Katti Sachin Katti on linkedin
    OpenAI's compute strategy requires both breadth and leverage: deep partnerships with the industry's leading accelerator, cloud, systems …
  • Richard Ho Richard Ho on linkedin
    🌶️ Jalapeño's first detailed performance results show industry-leading speed and efficiency for AI inference. …
  • @edzitron.com Ed Zitron on bluesky
    Even in this example running its own chip on its own infrastructure, OpenAI refuses to actually document how much *cheaper* it would be to use its own chip, and doesn't even bother to benchmark on its own frontier models, defaulting to open source.  Useless.  [embedded post]
  • r/accelerate r on reddit
    Jalapeño's first results show industry-leading speed and efficiency in AI inference
  • r/singularity r on reddit
    OpenAI blog post on their new custom inference chip