A detailed look at Jalapeño, OpenAI's ASIC developed with Broadcom in 16 months, which beat Nvidia, AMD, and Google chips on multiple top open-weight models
SemiAnalysis:
SemiAnalysis
Context & Ripple Effects
Jalapeño began as a deliberate hedge: sources told the Financial Times in September 2025 that OpenAI had committed $10B in orders for an internally-used Broadcom co-design, and reporting since then framed it as part of a split strategy — Nvidia for training, Broadcom for inference — with expected savings of 20%-30% against Nvidia's backlogged GPU pricing. After the chip's June 2026 unveiling, the open question was whether OpenAI's own silicon could actually compete rather than just save money.
First-order effects
OpenAI now has measured evidence — 1.5x-1.9x more work per watt and lower latency versus Nvidia across GPT-OSS, DeepSeek R1, and Kimi K2.5 — that its own ASIC beats merchant GPUs on its dominant workload, directly validating the internal-deployment bet made when it committed the $10B.
Broadcom gains the industry's most credible proof point for its custom-silicon model, while Nvidia, AMD, and Google all lose the 'custom chips can't match leading accelerators' argument on record.
Second-order effects
Other frontier labs facing the same inference bill are pushed toward comparable Broadcom-style co-designs, turning what was a cost play into a competitive necessity; Broadcom's custom ASIC pipeline becomes the default route for labs that don't want to build their own silicon teams.
The bottleneck shifts downstream: discussion already flags HBM4 sourcing — one account argues OpenAI's ramp is capped if it depends on a single supplier — so memory allocation, not accelerator design, becomes the constraint that determines whether Jalapeño scales to Rubin-class deployments.
Third-order effects
If the benchmark pattern holds across future revisions, inference splits structurally from training: Nvidia retains the general-purpose and training markets while high-volume inference consolidates around workload-specific silicon owned by the labs themselves, shrinking the merchant-GPU pool exactly where volume is largest.
Public benchmarks like InferenceX become the procurement battleground for AI compute, replacing vendor datasheets the way SPEC did for CPUs — which favors any player willing to publish measured results over those selling on brand.
The trend: Frontier labs are moving high-volume inference onto self-owned, workload-specific ASICs built through merchant foundry partners, converting Nvidia's largest customers into its direct competitors.
Well, Nvidia has to sell 10M chips to lots of companies next year I'm pretty sure OpenAI will be super satisfied if 25%-35% of the compute deployed by them uses Jalapeno and its successors in 2028 and beyond They don't need/have to scale like Nvidia/Rubin
One thing I can say is that OAI's Jalapeño can't scale to the same level as Rubin. If they source HBM4 exclusively from Samsung, there will be a limit to how far they can ramp. To scale beyond a certain point, they'll likely need to lower the HBM4 speed requirement so they can
OpenAI: “Today, we shared the first measured performance results from Jalapeño, OpenAI's first custom inference chip. On InferenceX, a public benchmark using GPT-OSS 120B, Jalapeño delivered more peak throughput per kilowatt and lower token latency than the commercial systems in
Jalapeno beats VR200 in A0 silicon with a vibe-ported, non-specdec implementation of DeepSeek 🌶️ 🌶️ And with much faster program execution than VR200, and with a B0 update landing imminently Congrats to the Jalapeno team!!!
No MTP, No PD disaggregation, Pure TP, still beats NVIDIA's Vera Rubin NVL72 on a third-party model, with A0 stepping. And B0 is 25% better. NVIDIA GPUs become HBM wrappers.
Data centers are power-limited and OpenAI's Jalapeño sets new frontiers for tokens per second per watt. Impressive and fast work from OpenAI's ASIC program.
OpenAI is designing for perf/W because they are currently limited by datacenter power envelopes (not budget or floor space) and have produced the leading accelerator in the industry by that metric It took 9 months to go from initial RTL to tapeout, and 6-8 more months for it to
> Using Codex with GPT-Astra, the team brought three open-weight models that were not part of Jalapeño's original production plan to high performance within two months from openai blog, two months is quite a lot here no? i'm a bit confuse by this
One of the most interesting parts of this is how the Jalapeno team (allegedly, and I think this is reasonable) used AI to accelerate the bringup of the chip on frontier(ish) models They can literally just profile the runs and hand it off to Codex for infinite iteration This
Since announcing Jalapeño, our first custom inference chip, we've been testing it and the system around it. The results show a major advance: more intelligence from every watt and faster responses, delivering both higher throughput and lower latency in one architecture without
Jalepeno is the best example we have so far of extremely high roi token spend by a large enterprise to create an entirely new business line. No surprise that it came first from OpenAI/Anthropic.