OpenAI Jalapeño vs Nvidia GB300: First Real Benchmarks
The Short Answer
At Hot Chips 2026, with results published August 25-26, 2026, OpenAI released the first real performance data for Jalapeño — the custom inference ASIC it co-developed with Broadcom.
The headline numbers, measured on SemiAnalysis’s InferenceX benchmark and run in OpenAI’s own lab alongside OpenAI engineers:
- 1.5x to 1.9x more throughput per kilowatt than Nvidia GB200 and GB300 rack systems
- 1.7x to 3.6x lower end-to-end latency
- 2.1x to 4.1x higher performance on the most interactive workloads
- 700W package power versus 1,400W for the GB300
The efficiency gap exists before performance is even factored in. Half the power draw is the starting point, not the result.
The Numbers Side by Side
| Jalapeño | Nvidia GB300 | |
|---|---|---|
| Type | Custom inference ASIC | Merchant GPU rack system |
| Package power | 700W | 1,400W |
| Throughput / kW | 1.5-1.9x advantage | baseline |
| End-to-end latency | 1.7-3.6x lower | baseline |
| Interactive workloads | 2.1-4.1x higher perf | baseline |
| Availability | OpenAI internal only | Purchasable |
| Workload | Inference only | Training + inference |
| Design partner | Broadcom | in-house |
Last verified: August 27, 2026. Benchmark: SemiAnalysis InferenceX, public methodology, executed in OpenAI’s lab.
Why OpenAI Optimised for Watts, Not FLOPS
This is the part most coverage skipped, and it is the part that actually explains the chip.
OpenAI is limited by datacenter power, not by budget or floorspace. When power is the binding constraint, the metric that governs how much product you can ship is tokens per megawatt — not tokens per dollar of silicon, and not peak FLOPS.
Jensen Huang made effectively the same argument at Computex 2026, citing perf/W, reliability and long lifetime as the axes that matter at hyperscale. The two companies agree on the metric. They disagree on who should build the chip that wins it.
An inference-only ASIC has a structural advantage here. It sheds everything a general-purpose training GPU must carry: training-specific datapaths, the flexibility to run arbitrary workloads, the ecosystem surface. If you know exactly which models you will serve and you control the whole stack, you can spend that transistor budget on the thing you actually do.
Three Reasons to Discount the Claim
Take the numbers seriously. Do not take them uncritically.
1. The comparison target is a generation behind. Jalapeño was benchmarked against GB200 and GB300. The efficiency advantage shrinks against Nvidia’s Vera Rubin generation. A 1.9x lead over the chip Nvidia is replacing is a different claim from a 1.9x lead over what Nvidia ships next.
2. It is a first-party benchmark, even with third-party involvement. SemiAnalysis engineers ran InferenceX and published a hands-on account — genuinely better than a vendor slide. But the runs happened in OpenAI’s lab, on OpenAI’s stack, alongside OpenAI’s engineers, on workloads OpenAI selected. That is favourable ground.
3. Silicon that benchmarks well is not silicon that ships. OpenAI plans small-volume deployment by the end of 2026 and broader rollout during 2027. Between a Hot Chips presentation and a serving fleet lie yield, packaging, HBM supply, and the software maturity Nvidia spent fifteen years accumulating. Analysts covering the announcement called the performance claims preliminary until the chip reaches large-scale manufacturing — that is the correct posture.
What It Means for Nvidia
Nvidia’s stock moved on the news, and Broadcom’s moved with it. The market read is roughly right: this is a margin story, not a volume story.
OpenAI is among the largest inference buyers on earth. A credible in-house alternative at that account does not remove Nvidia’s revenue overnight — it removes Nvidia’s ability to price as the only option at the accounts that set the reference price for everyone else.
What Nvidia still holds:
- Training. Jalapeño is inference-only. Frontier training remains Nvidia’s.
- CUDA. Fifteen years of software gravity that a custom ASIC does not inherit.
- Everyone who is not OpenAI. Custom silicon requires the volume to amortise a multi-hundred-million-dollar design. Perhaps five organisations globally clear that bar.
The pattern is now unmistakable across 2026: Google’s TPU line, Amazon’s Trainium, Anthropic’s Samsung 2nm effort, and now Jalapeño. Every hyperscaler with sufficient inference volume is building its own inference silicon. The merchant-GPU market keeps growing; its share of the largest buyers’ spend does not.
What It Means for You
Almost certainly nothing you can act on. You cannot buy Jalapeño.
What you may eventually see, in OpenAI’s own framing, is “faster responses, more responsive agents and more reliable access” as demand grows. Translated: lower latency and fewer capacity errors on the OpenAI API, sometime after volume deployment in 2027.
If you are modelling API prices, the second-order effect is the interesting one. Cheaper inference on OpenAI’s side creates room to cut prices without cutting margin. GPT-5.6 Sol already dropped from $5/$30 to $4/$20 on August 21, 2026, and GPT-5.6 Luna fell 80% on July 30. Custom silicon is one of the structural reasons that direction can continue — but do not budget on a price cut that has not been announced.
Sources
- SemiAnalysis — OpenAI Jalapeño: Better Than Nvidia Blackwell
- TechCrunch — OpenAI’s Jalapeño chip is built for fast inference at scale, benchmarks show
- Tom’s Hardware — OpenAI’s 700W Jalapeño ASIC outpaces 1,400W Nvidia flagship GPU
- CNBC — OpenAI’s Jalapeño AI chip brings new ‘threat’ to Nvidia margins