The OpenAI Jalapeño chip, the company’s first custom inference ASIC, delivered 1.5x to 1.9x more AI work per watt than Nvidia’s Blackwell systems in SemiAnalysis InferenceX tests published August 25, 2026. It draws 700W against GB300’s 1,400W and cut end-to-end latency by up to 3.6x. The catch: these are engineering samples. Volume deployment does not arrive until 2027.
What is the OpenAI Jalapeño chip?
The OpenAI Jalapeño chip is a custom inference accelerator co-developed with Broadcom and fabricated on TSMC’s N3P node. It is built to serve tokens, not train models. OpenAI published its first third-party benchmarks this week, and they are better than any first-generation silicon has a right to be.
The headline spec: 13.4 PFLOPS of MXFP4 compute at a 700W rating, paired with HBM4 running at 15.4 TB/s of bandwidth. In sustained operation the part draws under 550W, according to the benchmark data reported by ForkLog.
Nvidia’s GB200 rack unit pulls 1,200W. GB300 pulls 1,400W. Rubin sits between 900W and 1,150W. Jalapeño is doing its work in roughly half the power envelope.
The timeline is the real story
OpenAI started design in mid-2024 and handed the chip to the fab in November 2025. That is nine months from first design to manufacturing handoff, and 16 months to tape-out — a schedule that normally takes a silicon team two to three years.
OpenAI says its own models helped design the chip. That claim is unverifiable from the outside, but the calendar is not.
“Jalapeño can serve more AI work per unit of power, while also returning responses more quickly,” said Richard Ho, OpenAI’s head of hardware, in comments reported by TechCrunch.
How much faster is Jalapeño than Nvidia Blackwell?
Across three open-weight models, Jalapeño roughly doubled Nvidia’s tokens per second per kilowatt while cutting latency by 43% to 72%. The gap widens as models get larger. On DeepSeek R1 670B, Jalapeño returned a first response in 1.65 seconds against GB300’s 5.99 seconds.
Here are the SemiAnalysis InferenceX results as reported by ForkLog:
| Model | Jalapeño (mixed TPS/kW) | Nvidia system | Nvidia (mixed TPS/kW) | Jalapeño latency | Nvidia latency |
|---|---|---|---|---|---|
| GPT-OSS 120B | 85,448 | GB200 | 44,960 | 1.03s | 1.80s |
| DeepSeek R1 670B | 19,641 | GB300 | 11,781 | 1.65s | 5.99s |
| Kimi K2.5 1T | 18,195 | GB300 | 11,862 | 1.56s | 5.31s |
On single-user throughput, Jalapeño hit roughly 1,400 tokens per second on GPT-OSS 120B and over 700 tokens per second on DeepSeek R1 670B.
The aggregate claims are wider still: 1.7x to 3.6x lower end-to-end latency and 2.1x to 4.1x higher performance on interactive workloads, per The Decoder. At matched decoding speeds, The Decoder reported token-throughput-per-kilowatt advantages of 54x to 104x — a number that only makes sense in the narrow regime where GPU batching collapses.
What SemiAnalysis actually said
“Usually first generation chips aren’t competitive, but OpenAI is beating Nvidia Blackwell and even Rubin,” SemiAnalysis CEO Dylan Patel said, per The Decoder.
That is a strong endorsement from an analyst house that sells research to the same hyperscalers buying Nvidia racks. Take it seriously. Take it with salt.
Why does performance per watt decide who wins?
Because power, not silicon, is the binding constraint on AI buildouts in 2026. Data center operators are queuing for grid interconnects measured in years. If a chip does the same work at half the watts, the same substation serves twice the revenue.
That math is why custom ASICs keep appearing. Every watt saved on inference is a watt available for a paying customer, and inference is now the majority of frontier-lab compute spend.
OpenAI CFO Sarah Friar framed it in cost terms: custom chips give the company “greater control over inference costs” and let it match hardware to specific tasks. Friar also said the chip “complements” existing partnerships rather than replacing them — corporate language for we are still buying your GPUs, please keep taking our calls.
We covered the same power-and-memory squeeze from the supply side in our piece on the Nvidia AI server price hike, and the economics of fast inference in Cerebras vs Groq.
What does this do to Nvidia’s margins?
Nothing this quarter. Nvidia reported Q2 fiscal 2027 revenue of $96.22 billion on August 26, beating the $92.07 billion consensus, with data center revenue of $89.02 billion — up 117% year over year, according to 24/7 Wall St. EPS came in at $2.22 against a $2.09 estimate.
Guidance was louder than the beat. Nvidia guided Q3 to $108 billion plus or minus 2%, with non-GAAP gross margins near 74% and no China data center compute revenue assumed.
“AI has reached its inflection point. It’s doing useful work. Its tokens are productive and profitable. Now, compute is revenue,” CEO Jensen Huang said on the call.
Nvidia also disclosed supply commitments of $279 billion, largely for Vera Rubin memory. That is a company buying ahead, not one bracing for demand loss.
The threat is 2028, not 2026
Custom silicon does not eat Nvidia’s revenue. It eats Nvidia’s pricing power. A 74% gross margin exists because there is no substitute at scale. Jalapeño is the first credible substitute built by Nvidia’s single largest customer.
NVDA closed at $213.05 before the print, down 3.04% on the week and up 14.37% year to date, per 24/7 Wall St. The stock has fallen after four of its last five earnings reports despite beating consensus three quarters running.
Who wins and who loses financially?
Broadcom is the clearest winner. It gets ASIC design revenue, a marquee reference customer, and validation that its custom-silicon business can beat the merchant-GPU incumbent on a first attempt. Nvidia is the clearest medium-term loser, though the damage lands in 2028 pricing, not 2026 volume.
- Broadcom — books high-margin custom ASIC revenue and proves the model. We covered its financing appetite in the Broadcom AI debt deal.
- TSMC — wins either way. N3P wafers are N3P wafers, whether the logo says Nvidia or OpenAI.
- HBM suppliers — Jalapeño uses HBM4 at 15.4 TB/s. More custom chips means more high-bandwidth memory demand, not less.
- OpenAI — gains leverage in every future GPU negotiation, which may be worth more than the chip itself. Its Nvidia relationship already shifted once, as we noted when Nvidia cut its OpenAI data center guarantee.
- Nvidia — keeps the volume through 2027, then defends 74% margins against a credible in-house alternative.
- Second-tier inference clouds — squeezed hardest. They rent GPUs at market rates and cannot design their own.
What’s the catch with the Jalapeño benchmarks?
Three catches, and they matter. Jalapeño exists as engineering samples only. Rubin is already shipping to customers. And the benchmark set was chosen by the chip’s owner, run on three open-weight models, with two of Nvidia’s standard optimizations absent from the comparison.
The Decoder reported that Jalapeño lacks multi-token prediction and speculative decoding optimizations. Those are exactly the techniques that close latency gaps on GPUs. Adding them later helps Jalapeño; adding them to the comparison today would narrow the gap.
The models tested were GPT-OSS 120B, DeepSeek R1 670B and Kimi K2.5 1T. Larger current-generation models — DeepSeek V4 Pro, Kimi K3 — were not tested at all. Neither, notably, was any GPT-5-class OpenAI frontier model, which is the workload the chip actually has to serve.
And the deployment schedule is honest about itself: very small volumes at the end of 2026, meaningful volume in 2027. OpenAI says a second generation is in advanced development and a third is in design.
A chip that wins benchmarks in August 2026 must still win against whatever Nvidia ships in 2027. That is a different race.
Frequently asked questions
Is the OpenAI Jalapeño chip available to buy?
No. It is an internal accelerator for OpenAI’s own inference fleet, currently at engineering-sample stage. Small-volume deployment starts at the end of 2026, with wider rollout in 2027. There is no external sales channel announced.
Who manufactures the Jalapeño chip?
Broadcom co-developed it with OpenAI, and TSMC fabricates it on the N3P process node. The benchmarked silicon is B0 stepping, meaning at least one revision past first tape-out.
Does Jalapeño beat Nvidia’s Rubin?
On the perf-per-watt figures SemiAnalysis published, yes — 1.5x to 1.9x. But Rubin is shipping to paying customers now and Jalapeño is not, so the comparison is between a product and a prototype.
Can Jalapeño train models?
No. It is an inference-only design. OpenAI still needs GPUs for training, which is why CFO Sarah Friar described the chip as complementing rather than replacing existing supplier relationships.
How much power does Jalapeño use?
It is rated at 700W and reportedly sustains under 550W in operation. Nvidia’s GB200 draws 1,200W and GB300 draws 1,400W, so Jalapeño operates in roughly half the envelope.
Did Nvidia’s earnings show any damage from custom chips?
None yet. Data center revenue grew 117% year over year to $89.02 billion and Q3 guidance is $108 billion. Custom silicon is a 2028 margin question, not a 2026 revenue question.
What benchmark was used?
SemiAnalysis InferenceX, which measures mixed tokens per second per kilowatt alongside end-to-end latency. It is a third-party benchmark, but the model selection and test configuration came from the chip’s owner.
The bottom line
Jalapeño is the most serious first-generation AI accelerator anyone has produced, and the power numbers are the part that should worry Nvidia. Half the watts for double the tokens is not a rounding error; it is a structural argument for custom silicon at every lab large enough to fund a design team.
But the trade here is not “sell Nvidia.” Nvidia just printed $96.22 billion in a quarter and guided to $108 billion. The trade is that Nvidia’s 74% gross margin now has an expiry date attached, and the market will start pricing that date long before 2028 arrives.
The honest read: OpenAI has proven it can build a chip. It has not yet proven it can build ten million of them, on schedule, while Nvidia iterates annually. Benchmarks are cheap. Yield is not.
Sources
- TechCrunch — OpenAI’s Jalapeño chip is built for fast inference at scale, benchmarks show
- The Decoder — OpenAI’s first custom chip Jalapeño reportedly beats Nvidia’s Blackwell and Rubin
- ForkLog — OpenAI says Jalapeño outperforms Nvidia Blackwell in tests
- CNBC — OpenAI’s Jalapeño AI chip brings new threat to Nvidia margins
- 24/7 Wall St. — Nvidia Q2 2027: $96 Billion Quarter Fueled by 117% Data Center Surge