/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

A detailed look at Jalapeño, OpenAI's ASIC developed with Broadcom in 16 months, which beat Nvidia, AMD, and Google chips on multiple top open-weight models

SemiAnalysis:

SemiAnalysis

Context & Ripple Effects

Jalapeño began as a deliberate hedge: sources told the Financial Times in September 2025 that OpenAI had committed $10B in orders for an internally-used Broadcom co-design, and reporting since then framed it as part of a split strategy — Nvidia for training, Broadcom for inference — with expected savings of 20%-30% against Nvidia's backlogged GPU pricing. After the chip's June 2026 unveiling, the open question was whether OpenAI's own silicon could actually compete rather than just save money.

First-order effects

  • OpenAI now has measured evidence — 1.5x-1.9x more work per watt and lower latency versus Nvidia across GPT-OSS, DeepSeek R1, and Kimi K2.5 — that its own ASIC beats merchant GPUs on its dominant workload, directly validating the internal-deployment bet made when it committed the $10B.
  • Broadcom gains the industry's most credible proof point for its custom-silicon model, while Nvidia, AMD, and Google all lose the 'custom chips can't match leading accelerators' argument on record.

Second-order effects

  • Other frontier labs facing the same inference bill are pushed toward comparable Broadcom-style co-designs, turning what was a cost play into a competitive necessity; Broadcom's custom ASIC pipeline becomes the default route for labs that don't want to build their own silicon teams.
  • The bottleneck shifts downstream: discussion already flags HBM4 sourcing — one account argues OpenAI's ramp is capped if it depends on a single supplier — so memory allocation, not accelerator design, becomes the constraint that determines whether Jalapeño scales to Rubin-class deployments.

Third-order effects

  • If the benchmark pattern holds across future revisions, inference splits structurally from training: Nvidia retains the general-purpose and training markets while high-volume inference consolidates around workload-specific silicon owned by the labs themselves, shrinking the merchant-GPU pool exactly where volume is largest.
  • Public benchmarks like InferenceX become the procurement battleground for AI compute, replacing vendor datasheets the way SPEC did for CPUs — which favors any player willing to publish measured results over those selling on brand.

The trend: Frontier labs are moving high-volume inference onto self-owned, workload-specific ASICs built through merchant foundry partners, converting Nvidia's largest customers into its direct competitors.

Discussion

  • @zephyr_z9 @zephyr_z9 on x
    Well, Nvidia has to sell 10M chips to lots of companies next year I'm pretty sure OpenAI will be super satisfied if 25%-35% of the compute deployed by them uses Jalapeno and its successors in 2028 and beyond They don't need/have to scale like Nvidia/Rubin
  • @gdb Greg Brockman on x
    inference numbers published for jalapeno, team did an amazing job https://openai.com/...
  • @jukan05 Jukan on x
    One thing I can say is that OAI's Jalapeño can't scale to the same level as Rubin. If they source HBM4 exclusively from Samsung, there will be a limit to how far they can ramp. To scale beyond a certain point, they'll likely need to lower the HBM4 speed requirement so they can
  • @teortaxestex @teortaxestex on x
    Jensen really has no purely technical moat anymore, huh at least on inference, frontier labs can cut him out
  • @firstadopter Tae Kim on x
    OpenAI: “Today, we shared the first measured performance results from Jalapeño, OpenAI's first custom inference chip. On InferenceX, a public benchmark using GPT-OSS 120B, Jalapeño delivered more peak throughput per kilowatt and lower token latency than the commercial systems in
  • r/NVDA_Stock r on reddit
    OpenAI' Jalapeño: Better Than Nvidia Blackwell
  • @scaling01 @scaling01 on x
    it's more interesting that OpenAI had Astra for ~2 months
  • @itsclivetime Clive Chan on x
    Jalapeno beats VR200 in A0 silicon with a vibe-ported, non-specdec implementation of DeepSeek 🌶️ 🌶️ And with much faster program execution than VR200, and with a B0 update landing imminently Congrats to the Jalapeno team!!!
  • @youjiacheng You Jiacheng on x
    No MTP, No PD disaggregation, Pure TP, still beats NVIDIA's Vera Rubin NVL72 on a third-party model, with A0 stepping. And B0 is 25% better. NVIDIA GPUs become HBM wrappers.
  • @liamfedus Liam Fedus on x
    Data centers are power-limited and OpenAI's Jalapeño sets new frontiers for tokens per second per watt. Impressive and fast work from OpenAI's ASIC program.
  • @eladgil Elad Gil on x
    Impressive speed to tape out by OpenAI https://newsletter.semianalysis.com/ ...
  • @jordannanos Jordan Nanos on x
    OpenAI is designing for perf/W because they are currently limited by datacenter power envelopes (not budget or floor space) and have produced the leading accelerator in the industry by that metric It took 9 months to go from initial RTL to tapeout, and 6-8 more months for it to
  • @eliebakouch Elie on x
    > Using Codex with GPT-Astra, the team brought three open-weight models that were not part of Jalapeño's original production plan to high performance within two months from openai blog, two months is quite a lot here no? i'm a bit confuse by this
  • @eliebakouch Elie on x
    1. openai doesn't use MLA for attention mechanism 2. they implemented MLA kernel with codex without any human supervision
  • @noslawextratost Cam Quilici on x
    One of the most interesting parts of this is how the Jalapeno team (allegedly, and I think this is reasonable) used AI to accelerate the bringup of the chip on frontier(ish) models They can literally just profile the runs and hand it off to Codex for infinite iteration This
  • @thelokasiffers Giulio on x
    maybe I'm retarded but isn't it expected for an ASIC to outperform a general-purpose GPU on the workload the ASIC was specifically designed for?
  • @cdleary @cdleary on x
    it's a good chip ser agree w/ the article: the team is cracked and delivered like you wouldn't believe (gonna be a lot of cope) the spice must flow 🌶️
  • @semianalysis_ @semianalysis_ on x
    OpenAI Jalapeño: Better Than Nvidia Blackwell OpenAI's self-designed ASIC compared with Rubin, Jalapeño's TCO, throughput per MW, and spicy deets https://newsletter.semianalysis.com/ ...
  • Madeleine Laitz, PhD Madeleine Laitz, PhD on linkedin
    Really excited to see Jalapeño out in the world from OpenAI.  The most interesting part to me is the strategy …
  • r/OpenAI r on reddit
    OpenAI Jalapeño: Better Than Nvidia Blackwell
  • r/AMD_Stock r on reddit
    OpenAI' Jalapeño: Better Than Nvidia Blackwell
  • @openai @openai on x
    Since announcing Jalapeño, our first custom inference chip, we've been testing it and the system around it. The results show a major advance: more intelligence from every watt and faster responses, delivering both higher throughput and lower latency in one architecture without
  • @maxkan Max Kan on x
    Jalepeno is the best example we have so far of extremely high roi token spend by a large enterprise to create an entirely new business line. No surprise that it came first from OpenAI/Anthropic.