Tag: Semiconductors

  • OpenAI Jalapeño Chip Beats Blackwell 1.9x Per Watt — Ships 2027

    The OpenAI Jalapeño chip, the company’s first custom inference ASIC, delivered 1.5x to 1.9x more AI work per watt than Nvidia’s Blackwell systems in SemiAnalysis InferenceX tests published August 25, 2026. It draws 700W against GB300’s 1,400W and cut end-to-end latency by up to 3.6x. The catch: these are engineering samples. Volume deployment does not arrive until 2027.

    What is the OpenAI Jalapeño chip?

    The OpenAI Jalapeño chip is a custom inference accelerator co-developed with Broadcom and fabricated on TSMC’s N3P node. It is built to serve tokens, not train models. OpenAI published its first third-party benchmarks this week, and they are better than any first-generation silicon has a right to be.

    The headline spec: 13.4 PFLOPS of MXFP4 compute at a 700W rating, paired with HBM4 running at 15.4 TB/s of bandwidth. In sustained operation the part draws under 550W, according to the benchmark data reported by ForkLog.

    Nvidia’s GB200 rack unit pulls 1,200W. GB300 pulls 1,400W. Rubin sits between 900W and 1,150W. Jalapeño is doing its work in roughly half the power envelope.

    The timeline is the real story

    OpenAI started design in mid-2024 and handed the chip to the fab in November 2025. That is nine months from first design to manufacturing handoff, and 16 months to tape-out — a schedule that normally takes a silicon team two to three years.

    OpenAI says its own models helped design the chip. That claim is unverifiable from the outside, but the calendar is not.

    “Jalapeño can serve more AI work per unit of power, while also returning responses more quickly,” said Richard Ho, OpenAI’s head of hardware, in comments reported by TechCrunch.

    How much faster is Jalapeño than Nvidia Blackwell?

    Across three open-weight models, Jalapeño roughly doubled Nvidia’s tokens per second per kilowatt while cutting latency by 43% to 72%. The gap widens as models get larger. On DeepSeek R1 670B, Jalapeño returned a first response in 1.65 seconds against GB300’s 5.99 seconds.

    Here are the SemiAnalysis InferenceX results as reported by ForkLog:

    Model Jalapeño (mixed TPS/kW) Nvidia system Nvidia (mixed TPS/kW) Jalapeño latency Nvidia latency
    GPT-OSS 120B 85,448 GB200 44,960 1.03s 1.80s
    DeepSeek R1 670B 19,641 GB300 11,781 1.65s 5.99s
    Kimi K2.5 1T 18,195 GB300 11,862 1.56s 5.31s

    On single-user throughput, Jalapeño hit roughly 1,400 tokens per second on GPT-OSS 120B and over 700 tokens per second on DeepSeek R1 670B.

    The aggregate claims are wider still: 1.7x to 3.6x lower end-to-end latency and 2.1x to 4.1x higher performance on interactive workloads, per The Decoder. At matched decoding speeds, The Decoder reported token-throughput-per-kilowatt advantages of 54x to 104x — a number that only makes sense in the narrow regime where GPU batching collapses.

    What SemiAnalysis actually said

    “Usually first generation chips aren’t competitive, but OpenAI is beating Nvidia Blackwell and even Rubin,” SemiAnalysis CEO Dylan Patel said, per The Decoder.

    That is a strong endorsement from an analyst house that sells research to the same hyperscalers buying Nvidia racks. Take it seriously. Take it with salt.

    Why does performance per watt decide who wins?

    Because power, not silicon, is the binding constraint on AI buildouts in 2026. Data center operators are queuing for grid interconnects measured in years. If a chip does the same work at half the watts, the same substation serves twice the revenue.

    That math is why custom ASICs keep appearing. Every watt saved on inference is a watt available for a paying customer, and inference is now the majority of frontier-lab compute spend.

    OpenAI CFO Sarah Friar framed it in cost terms: custom chips give the company “greater control over inference costs” and let it match hardware to specific tasks. Friar also said the chip “complements” existing partnerships rather than replacing them — corporate language for we are still buying your GPUs, please keep taking our calls.

    We covered the same power-and-memory squeeze from the supply side in our piece on the Nvidia AI server price hike, and the economics of fast inference in Cerebras vs Groq.

    What does this do to Nvidia’s margins?

    Nothing this quarter. Nvidia reported Q2 fiscal 2027 revenue of $96.22 billion on August 26, beating the $92.07 billion consensus, with data center revenue of $89.02 billion — up 117% year over year, according to 24/7 Wall St. EPS came in at $2.22 against a $2.09 estimate.

    Guidance was louder than the beat. Nvidia guided Q3 to $108 billion plus or minus 2%, with non-GAAP gross margins near 74% and no China data center compute revenue assumed.

    “AI has reached its inflection point. It’s doing useful work. Its tokens are productive and profitable. Now, compute is revenue,” CEO Jensen Huang said on the call.

    Nvidia also disclosed supply commitments of $279 billion, largely for Vera Rubin memory. That is a company buying ahead, not one bracing for demand loss.

    The threat is 2028, not 2026

    Custom silicon does not eat Nvidia’s revenue. It eats Nvidia’s pricing power. A 74% gross margin exists because there is no substitute at scale. Jalapeño is the first credible substitute built by Nvidia’s single largest customer.

    NVDA closed at $213.05 before the print, down 3.04% on the week and up 14.37% year to date, per 24/7 Wall St. The stock has fallen after four of its last five earnings reports despite beating consensus three quarters running.

    Who wins and who loses financially?

    Broadcom is the clearest winner. It gets ASIC design revenue, a marquee reference customer, and validation that its custom-silicon business can beat the merchant-GPU incumbent on a first attempt. Nvidia is the clearest medium-term loser, though the damage lands in 2028 pricing, not 2026 volume.

    • Broadcom — books high-margin custom ASIC revenue and proves the model. We covered its financing appetite in the Broadcom AI debt deal.
    • TSMC — wins either way. N3P wafers are N3P wafers, whether the logo says Nvidia or OpenAI.
    • HBM suppliers — Jalapeño uses HBM4 at 15.4 TB/s. More custom chips means more high-bandwidth memory demand, not less.
    • OpenAI — gains leverage in every future GPU negotiation, which may be worth more than the chip itself. Its Nvidia relationship already shifted once, as we noted when Nvidia cut its OpenAI data center guarantee.
    • Nvidia — keeps the volume through 2027, then defends 74% margins against a credible in-house alternative.
    • Second-tier inference clouds — squeezed hardest. They rent GPUs at market rates and cannot design their own.

    What’s the catch with the Jalapeño benchmarks?

    Three catches, and they matter. Jalapeño exists as engineering samples only. Rubin is already shipping to customers. And the benchmark set was chosen by the chip’s owner, run on three open-weight models, with two of Nvidia’s standard optimizations absent from the comparison.

    The Decoder reported that Jalapeño lacks multi-token prediction and speculative decoding optimizations. Those are exactly the techniques that close latency gaps on GPUs. Adding them later helps Jalapeño; adding them to the comparison today would narrow the gap.

    The models tested were GPT-OSS 120B, DeepSeek R1 670B and Kimi K2.5 1T. Larger current-generation models — DeepSeek V4 Pro, Kimi K3 — were not tested at all. Neither, notably, was any GPT-5-class OpenAI frontier model, which is the workload the chip actually has to serve.

    And the deployment schedule is honest about itself: very small volumes at the end of 2026, meaningful volume in 2027. OpenAI says a second generation is in advanced development and a third is in design.

    A chip that wins benchmarks in August 2026 must still win against whatever Nvidia ships in 2027. That is a different race.

    Frequently asked questions

    Is the OpenAI Jalapeño chip available to buy?

    No. It is an internal accelerator for OpenAI’s own inference fleet, currently at engineering-sample stage. Small-volume deployment starts at the end of 2026, with wider rollout in 2027. There is no external sales channel announced.

    Who manufactures the Jalapeño chip?

    Broadcom co-developed it with OpenAI, and TSMC fabricates it on the N3P process node. The benchmarked silicon is B0 stepping, meaning at least one revision past first tape-out.

    Does Jalapeño beat Nvidia’s Rubin?

    On the perf-per-watt figures SemiAnalysis published, yes — 1.5x to 1.9x. But Rubin is shipping to paying customers now and Jalapeño is not, so the comparison is between a product and a prototype.

    Can Jalapeño train models?

    No. It is an inference-only design. OpenAI still needs GPUs for training, which is why CFO Sarah Friar described the chip as complementing rather than replacing existing supplier relationships.

    How much power does Jalapeño use?

    It is rated at 700W and reportedly sustains under 550W in operation. Nvidia’s GB200 draws 1,200W and GB300 draws 1,400W, so Jalapeño operates in roughly half the envelope.

    Did Nvidia’s earnings show any damage from custom chips?

    None yet. Data center revenue grew 117% year over year to $89.02 billion and Q3 guidance is $108 billion. Custom silicon is a 2028 margin question, not a 2026 revenue question.

    What benchmark was used?

    SemiAnalysis InferenceX, which measures mixed tokens per second per kilowatt alongside end-to-end latency. It is a third-party benchmark, but the model selection and test configuration came from the chip’s owner.

    The bottom line

    Jalapeño is the most serious first-generation AI accelerator anyone has produced, and the power numbers are the part that should worry Nvidia. Half the watts for double the tokens is not a rounding error; it is a structural argument for custom silicon at every lab large enough to fund a design team.

    But the trade here is not “sell Nvidia.” Nvidia just printed $96.22 billion in a quarter and guided to $108 billion. The trade is that Nvidia’s 74% gross margin now has an expiry date attached, and the market will start pricing that date long before 2028 arrives.

    The honest read: OpenAI has proven it can build a chip. It has not yet proven it can build ten million of them, on schedule, while Nvidia iterates annually. Benchmarks are cheap. Yield is not.

    Sources

  • Nvidia AI Server Price Hike: 15% More, and Memory Is Why

    Nvidia is raising AI server prices by more than 15% on Grace Blackwell and Vera Rubin systems shipping in early 2027, Bloomberg reported on August 24, 2026. Memory is the reason. UBS puts memory at 62% of a Vera Rubin superchip’s $38,902 bill of materials, up from 53% on Grace Blackwell. TrendForce estimates the hike adds at least $5 billion to a 1-gigawatt data center.

    For three years the AI trade had one simple rule: Nvidia sets the price, and everyone pays it. That rule still holds. What changed is who Nvidia is paying.

    The Nvidia AI server price hike is not a margin grab. It is a pass-through. And the numbers underneath it say the memory makers, not the GPU designer, now control the cost curve of the AI build-out.

    How much is Nvidia raising AI server prices?

    More than 15% in many cases, effective on systems shipped early next year. Bloomberg reported the increases on August 24, citing people familiar with the matter. TrendForce, summarizing the same reporting, said some configurations could reach 17%. Nvidia did not respond to requests for comment.

    The warnings did not go to the cloud giants directly. According to Bloomberg, Nvidia notified the contract server manufacturers that assemble systems for Microsoft, Alphabet’s Google, and Oracle.

    That routing matters. The ODMs absorb the notice first, then reprice their own quotes. The cloud buyers find out when the invoice changes.

    Which systems are affected

    • Grace Blackwell systems — the current generation, still shipping in volume.
    • Vera Rubin systems — the next generation, with first shipments in early 2027.
    • Increases vary by chip generation and by memory configuration, per Bloomberg. Denser memory builds take the larger hit.

    Why is Nvidia raising prices now?

    Because memory has gone from a line item to the line item. Morgan Stanley estimates GPU silicon has fallen from more than 80% of AI server cost to roughly half that level in next-generation systems. The gap did not close because GPUs got cheaper. It closed because DRAM got expensive.

    A Vera Rubin NVL72 rack carries 74.7 TB of DRAM — 20.7 TB of HBM4 plus 54 TB of LPDDR5X, according to UBS’s teardown. That is the DRAM content of roughly 4,500 smartphones in a single rack.

    Every one of those bits is bought in the tightest memory market in a decade.

    What UBS found inside a Vera Rubin superchip

    UBS’s bill-of-materials analysis is the clearest picture available of where the money actually goes.

    Component Cost Share of superchip
    Total Vera Rubin superchip $38,902 100%
    All memory $24,297 62%
    SOCAMM2 LPDDR5X $19,355 49.8%
    HBM4 $4,943 12.7%
    Everything else $14,605 38%

    On Grace Blackwell, UBS put memory at 53% of cost. On Vera Rubin it is 62%, and the absolute memory bill rose about 2.5x between generations.

    One caveat worth holding onto: that 2.5x blends two different things. Vera Rubin carries more memory and pays more per gigabyte. It is not a pure price signal.

    How much has DRAM actually gone up?

    Steeply, and for longer than most forecasts allowed. TrendForce data cited by Tom’s Hardware shows conventional DRAM contract prices rising 90–95% quarter-over-quarter in Q1 2026 and a projected 58–63% in Q2 2026. Server DRAM is expected to climb every quarter through the second half of 2027.

    The consumer market tells the same story in plainer numbers. A mainstream 32GB DDR5-6000 kit runs about $392 today against $110–$140 a year ago, per Tom’s Hardware.

    Supply was committed early. SK hynix had sold out its entire 2026 production capacity by October 2025. Samsung and SK hynix raised 2026 HBM3E prices by roughly 20%.

    And HBM makes the squeeze worse mechanically: it consumes roughly four times the wafer area of conventional DRAM per bit shipped. Every HBM4 order crowds out ordinary server memory on the same fab.

    Who profits from the Nvidia AI server price hike?

    Not Nvidia, on the arithmetic. The memory suppliers capture the increase, the ODMs pass it through, and the hyperscalers eat it. Nvidia’s role here is closer to toll collector than beneficiary — and its own gross margin may be the quiet casualty.

    Work the math. Nvidia runs roughly a 75% gross margin, so the bill of materials is about 25% of the sale price. If memory is 62% of that BOM, memory is about 15.5% of the price. A 2.5x memory cost increase adds roughly 23 points of price to cost.

    A 15% price hike does not cover 23 points. Something has to give.

    Three readings, and the market has not settled on one:

    1. The 15% is an opening installment. More increases follow as 2027 contracts reprice.
    2. Nvidia is absorbing the difference. Gross margin drifts from ~75% toward the high 60s.
    3. The 2.5x is generational, not inflationary. Higher memory content is sold at a higher system ASP, so the comparison overstates the pass-through problem.

    Reading three is the most likely and the least discussed. It is also the one that would let Nvidia keep its margin story intact — which is precisely why it deserves scrutiny rather than acceptance.

    Why this matters for AI capex

    Because it reprices the entire build-out. TrendForce estimates the increase adds at least $5 billion to the cost of a 1-gigawatt AI data center. At the scale hyperscalers are now committing to, that is not a rounding error — it is a line in the capital plan that did not exist last quarter.

    The second-order effects are where this gets interesting.

    The uncomfortable version: AI compute has been getting cheaper per unit of intelligence for three straight years. This is the first credible input cost that pushes the other way.

    This post is reporting and analysis, not financial advice.

    Frequently asked questions

    How much is Nvidia raising AI server prices?

    More than 15% in many cases, with some configurations reaching 17% per TrendForce. Increases vary by chip generation and memory configuration.

    When do the new prices take effect?

    On systems shipped in early 2027, according to Bloomberg’s August 24, 2026 report.

    Which Nvidia systems are affected?

    Grace Blackwell and Vera Rubin server systems. Both are rack-scale platforms sold to cloud and enterprise data center operators.

    Why are AI server prices going up?

    Memory costs. UBS puts memory at 62% of a Vera Rubin superchip’s cost, and DRAM contract prices have risen every quarter through 2026 amid an HBM-driven supply squeeze.

    Who was notified about the price increases?

    Contract server manufacturers that build systems for Microsoft, Google, and Oracle, per Bloomberg. Nvidia did not comment publicly.

    How much does this add to a data center?

    TrendForce estimates at least $5 billion in additional cost for a 1-gigawatt AI data center.

    Does this hurt Nvidia’s margins?

    Possibly. Nvidia runs roughly a 75% gross margin. If memory costs rose 2.5x generationally, a 15% price increase may not fully offset it — though part of that increase reflects more memory content per system, not pure inflation.

    The bottom line

    The Nvidia AI server price hike is the clearest sign yet that the AI supply chain’s power center is shifting. For three years the scarce input was GPU wafer allocation. In 2027 it is memory, and the companies that own it — SK hynix, Samsung, Micron — are the ones setting terms.

    Watch two things next. First, whether Nvidia’s gross margin guidance holds through the fiscal year, because that is where the pass-through gap shows up. Second, whether any hyperscaler publicly revises a gigawatt commitment. The first cost-driven downgrade of an announced buildout would tell you the memory squeeze has stopped being an engineering problem and started being a financial one.

    Sources

  • Google Marvell Chip Deal: $12.2B Warrant, $120B Catch

    Google secured a warrant for 58,970,907 Marvell shares at $206.58 each — about $12.2 billion — under a custom silicon agreement disclosed on August 19, 2026. Marvell’s 8-K shows only 1,360,867 shares vest on time. The rest unlock in 240 tranches, one per $500 million of custom product revenue: $120 billion of chip purchases through fiscal 2033.

    The Google Marvell chip deal is the clearest sign yet that hyperscalers no longer just buy silicon. They take equity in the companies that build it.

    Marvell Technology stock jumped 13% on the disclosure. Broadcom fell 3%. Alphabet did not move at all.

    What is the Google Marvell chip deal?

    It is a custom silicon supply agreement signed July 29, 2026, paired with a stock warrant issued August 18, 2026. Marvell will design chips across five categories for Google’s TPU infrastructure. In exchange, Google holds an option on roughly 7% of Marvell, priced today and payable later.

    According to Marvell’s 8-K filing with the SEC, the warrant expires August 18, 2033.

    The five product lines Marvell will supply, per analysis from The Futurum Group:

    • Inference accelerators
    • Storage controllers
    • Network interface controllers
    • Memory interface controllers
    • Near-memory compute

    That is not one chip. That is a seat at every layer of the rack.

    How much is the warrant actually worth?

    At the $206.58 strike price, full exercise costs Google about $12.18 billion and delivers 58,970,907 shares. But the headline number is a ceiling, not a payment. Google owes nothing today. Almost the entire position is contingent on purchase volume Marvell has never come close to booking from a single customer.

    Term Detail
    Warrant shares 58,970,907
    Exercise price $206.58 per share
    Value at full exercise ~$12.18 billion
    Time-based tranche 1,360,867 shares, equal quarterly installments in year one
    Performance tranches 240 tranches, one per $500M of custom product revenue
    Implied purchase total $120 billion
    Vesting window Q3 fiscal 2027 through end of fiscal 2033
    Expiration August 18, 2033
    Commercial agreement signed July 29, 2026

    The vesting math nobody put in the headline

    Divide 240 tranches by the roughly six and a half years between Q3 fiscal 2027 and the end of fiscal 2033. Futurum calculates Google would need to average close to $18 billion a year in custom purchases from Marvell to unlock the full warrant.

    Hold that number. It matters in a moment.

    Why would Google take equity in its own supplier?

    Because it converts a procurement line into an asset. If Google spends $120 billion with Marvell and Marvell’s stock rises on that revenue, Google captures part of the gain it created. If Google spends nothing, the warrant lapses and costs it nothing.

    The structure is asymmetric by design. Google pays with optionality, not cash.

    It also locks Marvell in. A supplier whose largest shareholder-in-waiting is its largest customer has limited leverage on price. That is the quiet half of the deal.

    Variations of this circular financing keep appearing across the sector — most visibly when Nvidia cut its OpenAI data center guarantee from $250B to $120B, and again in Broadcom’s up-to-$100 billion debt raise to fund Anthropic chips.

    Does Marvell replace Broadcom as Google’s TPU partner?

    No. Broadcom remains Google’s primary TPU design partner under a long-term agreement running through 2031. Morningstar analyst William Kerwin, quoted by TheStreet, called the deal “a strong win for Marvell” while noting Google was “adding new suppliers rather than dropping Broadcom.”

    The read is capacity, not replacement. Broadcom’s design teams are booked on core accelerator generations. Marvell picks up memory expansion, decode-focused inference, and interconnect controllers.

    Broadcom’s 3% drop on the news looks like a market pricing in a smaller share of a much larger pie.

    Can Marvell realistically deliver $120 billion?

    This is where the number starts to strain. Marvell’s Q1 fiscal 2027 results show total net revenue of $2.418 billion for the quarter ended May 2, 2026, with data center at $1.833 billion — 76% of the business and up 28% year over year.

    Guidance for Q2 is $2.700 billion, plus or minus 5%. Annualize that and Marvell is a roughly $10.8 billion revenue company.

    Now compare. To fully vest the warrant, Google alone would need to buy about $18 billion of custom silicon a year — roughly 1.7 times everything Marvell currently sells to every customer combined.

    Management’s own stated target is more than $10 billion in custom revenue by fiscal 2029, across all customers. The Google ceiling sits an order of magnitude above the plan.

    Treat $120 billion as a theoretical maximum with a marketing function, not a forecast. The tranche structure exists precisely because neither side expects the top of the range.

    How did the market react?

    Sharply, and selectively. On August 19, 2026, Marvell rose 13% to $243.66 while Broadcom fell 3% to $369.13 and Alphabet closed unchanged at $342.67, according to 24/7 Wall St.

    Alphabet’s flat tape is the most interesting line in that table. A $120 billion purchase commitment moved the buyer’s stock zero percent.

    That tells you the market already assumed Google would spend the money somewhere. Only the recipient was in question.

    Dilution is real but modest: full exercise cuts existing shareholders by roughly 6.3% to 6.7% and would make Google approximately Marvell’s fifth-largest investor, per TheStreet.

    Why this matters

    Custom silicon is where the AI infrastructure margin is migrating. Every hyperscaler that designs its own accelerator takes revenue that would otherwise flow to Nvidia — and hands part of it to a merchant design partner like Broadcom or Marvell.

    The warrant structure is the new template. Compute buyers are increasingly paid in equity for their own demand. That is what Nvidia’s $6 billion Poolside arrangement did in software, and it is the same logic investors are pricing into custom-inference startups like Etched at a $21 billion valuation.

    For investors, the practical question is not whether the $120 billion lands. It is whether Marvell’s custom design wins convert into recognized revenue on the quarterly cadence the tranches imply. Watch the custom line, not the headline.

    This post is reporting and analysis, not financial advice.

    Frequently asked questions

    How many Marvell shares does the Google warrant cover?

    58,970,907 shares at an exercise price of $206.58, worth about $12.18 billion at full exercise, per Marvell’s 8-K.

    When does the Google Marvell warrant expire?

    August 18, 2033. Vesting runs from the third quarter of fiscal 2027 through the end of fiscal 2033.

    What has to happen for the full warrant to vest?

    Beyond 1,360,867 time-based shares, tranches vest one at a time for every $500 million of custom product revenue — 240 tranches, or $120 billion total.

    Is Google dropping Broadcom for Marvell?

    No. Broadcom holds a TPU design agreement through 2031 and remains the primary partner. Marvell is being added across adjacent chip categories.

    How much dilution do Marvell shareholders face?

    Roughly 6.3% to 6.7% if the warrant is fully exercised, which would put Google around fifth among Marvell’s largest holders.

    What is Marvell’s current revenue?

    $2.418 billion in the quarter ended May 2, 2026, with Q2 guidance of $2.700 billion plus or minus 5%.

    Did Alphabet stock move on the news?

    No. Alphabet closed unchanged at $342.67 on August 19, 2026, while Marvell rose 13% and Broadcom fell 3%.

    The bottom line

    The Google Marvell chip deal is a genuine design win wrapped in a number that will not be met. Marvell gets multi-year attachment across five product categories inside the largest custom accelerator program outside Nvidia. Google gets a free option on the value it creates by spending.

    The next real datapoint is Marvell’s custom product revenue line. Each $500 million tranche is a public scoreboard — a rare case of customer concentration disclosed quarter by quarter through a vesting schedule.

    If two or three tranches clear in fiscal 2028, the thesis holds. If the line stays flat while the stock trades on $120 billion, the gap closes the hard way.

    Sources