Tag: Nvidia

  • Nvidia Hugging Face Acquisition: $12.9 Billion at 80x Revenue

    The Nvidia Hugging Face acquisition values the open-source model hub at $12.9 billion, according to The Information — roughly 80 times its ~$150 million in annualized revenue. Hugging Face turned down a $500 million Nvidia investment at a $7 billion valuation less than a year ago. Neither company has confirmed the deal. It would be Nvidia’s second-largest acquisition ever, behind the $20 billion Groq purchase.

    How much is Nvidia paying for Hugging Face?

    Nvidia has agreed to pay approximately $12.9 billion for Hugging Face, The Information reported on August 27, citing a person familiar with the transaction. The deal is agreed but not signed. It can still collapse.

    CNBC and Fortune both matched the report the same day. Business Insider first reported Nvidia’s takeover interest.

    Neither Nvidia nor Hugging Face responded to requests for comment, per Quartz. That silence matters — nothing here is a signed, disclosed transaction yet.

    What Hugging Face was worth before

    Hugging Face last priced itself at $4.5 billion in a 2023 round led by Salesforce Ventures, with Alphabet’s GV, IBM Ventures and — notably — Nvidia participating, according to TechCrunch.

    In late 2025, Nvidia offered $500 million at a $7 billion valuation. Hugging Face said no. TechCrunch reports the company declined because it did not want a single dominant investor.

    Roughly nine months later, it is selling outright to that same investor for nearly double the valuation it rejected.

    Date Event Valuation
    2023 Series funding led by Salesforce Ventures $4.5 billion
    Late 2025 $500M Nvidia investment offer — declined $7 billion
    Aug 27, 2026 Reported acquisition agreement $12.9 billion

    Why is Nvidia buying an open-source model hub?

    Nvidia is buying distribution, not revenue. Hugging Face is where developers publish, discover and download open models — Tom’s Hardware calls it a “GitHub-like repository” for AI. Owning the shelf is worth more to Nvidia than the $150 million the shelf currently earns.

    The strategic timing is not subtle. Nvidia’s largest customers are building silicon that competes with its own.

    The lock-in play

    Hugging Face’s Inference Endpoints today support AWS Inferentia, AMD Instinct, Google TPU, Intel CPUs and Nvidia accelerators, per Tom’s Hardware. It is deliberately vendor-neutral.

    Under Nvidia, that neutrality is the first thing analysts expect to erode. If the default deployment path for every popular open model points at CUDA, Nvidia defends its installed base at the exact layer where switching decisions get made.

    That threat is real. OpenAI, Google, Amazon and Anthropic are all shipping or funding custom accelerators — see our coverage of OpenAI’s Jalapeño chip and the Broadcom debt package funding Anthropic’s silicon.

    The cloud re-entry play

    Nvidia scaled back DGX Cloud roughly a year ago, TechCrunch notes. Hugging Face gives it a consumer-facing compute surface again — and somewhere to route the capacity Nvidia has committed to but not fully sold.

    That is the least-discussed part of the rationale and possibly the most financially concrete one.

    Is $12.9 billion too much for $150 million in revenue?

    On the numbers, yes — by any conventional standard. Tom’s Hardware puts the deal at roughly 80 times forward revenue. Software acquisitions at 15–20x are already considered rich.

    Hugging Face’s revenue is growing fast. TechCrunch reports it moved from about $100 million to about $150 million annualized in roughly two months, and Tom’s Hardware says paying subscribers doubled in the first half of 2026. CEO Clem Delangue told TechCrunch last month the company was “close to profitability.”

    Here is the skeptical read. Nvidia is paying a strategic premium for neutrality it intends to end. The moment developers believe Hugging Face is a CUDA storefront rather than a Switzerland, some of them leave — and the asset Nvidia bought is worth less than the asset it paid for. AMD, Google and the open-weights community have every incentive to fund an alternative registry.

    Ten-year-old infrastructure businesses with $150 million in revenue do not usually command $12.9 billion. They command it when the buyer is defending a franchise.

    How does this fit Nvidia’s acquisition spree?

    It is the second-largest deal Nvidia has ever done, and the third multi-billion-dollar AI purchase in nine months. Nvidia has stopped behaving like a component supplier and started behaving like a platform consolidator.

    Target Reported price Announced What it buys
    Groq ~$20 billion Dec 2025 Inference architecture (LPU)
    Hugging Face $12.9 billion Aug 2026 Open-model distribution
    Poolside ~$6 billion Aug 2026 Model training capability

    The Groq deal — about $20 billion, reported by CNBC in December 2025 — was Nvidia’s largest on record. We covered the $6 billion Poolside purchase earlier this month.

    Nvidia can afford all of it in cash. Its Q2 fiscal 2027 results, for the quarter ended July 26, 2026, show $96.2 billion in revenue, up 106% year over year, and $59.7 billion in GAAP net income. Cash, marketable debt and marketable equity securities totaled roughly $99.3 billion.

    Why this matters

    Three things follow from this deal, and none of them are about Hugging Face.

    • The competitive threat is now priced. Nvidia is spending real money to defend against customers building their own chips. That is an admission the threat is material.
    • Open-source AI just got an owner. The default distribution point for open models moves inside a hardware vendor. Expect immediate pressure for a neutral alternative.
    • Strategic multiples are back. Eighty times revenue is a 2021-style number appearing in 2026, funded by operating cash rather than cheap debt.

    For investors, the read is about defensive capital allocation. Nvidia guided to about $108 billion for the current quarter — excluding any China data center compute revenue. A company growing that fast does not spend $12.9 billion on a $150 million business unless it sees a hole in the moat.

    It also fits a pattern of Nvidia paying to control adjacent chokepoints, much as Stripe paid over $7 billion for OpenRouter to sit on the AI token toll road. Note, too, that Nvidia recently cut its OpenAI data center guarantee from $250 billion to $120 billion — capital is being redirected, not simply added.

    This post is reporting and analysis, not financial advice.

    Frequently asked questions

    Is the Nvidia Hugging Face acquisition confirmed?

    No. The Information reported an agreement on August 27, 2026, and CNBC, Fortune and TechCrunch matched it. Neither company has commented publicly, and the deal is not signed.

    How much revenue does Hugging Face generate?

    Roughly $150 million annualized, up from about $100 million two months earlier, according to TechCrunch. The $12.9 billion price is about 80 times that figure.

    Why did Hugging Face reject Nvidia before?

    It declined a $500 million investment at a $7 billion valuation in late 2025 because it did not want a single dominant investor, TechCrunch reported.

    Will Hugging Face still support AMD and Google chips?

    Unknown. Its Inference Endpoints currently support AWS Inferentia, AMD Instinct, Google TPU, Intel CPUs and Nvidia accelerators. Nvidia has not said whether that continues.

    Is this Nvidia’s biggest acquisition?

    No. The roughly $20 billion Groq deal announced in December 2025 remains its largest, per CNBC. Hugging Face would rank second.

    Could regulators block the deal?

    No formal review has been reported. Antitrust scrutiny is plausible given Nvidia’s accelerator share and the platform’s role in model distribution, but nothing has been filed publicly.

    Can Nvidia pay cash?

    Comfortably. It reported $59.7 billion in GAAP net income in a single quarter and about $99.3 billion in cash and marketable securities as of July 26, 2026.

    The bottom line

    Nvidia is paying roughly 80 times revenue to own the front door of open-source AI. The financial case is thin; the defensive case is obvious.

    Watch three things next: whether the deal is actually signed, whether Hugging Face keeps supporting rival accelerators, and how quickly a neutral competitor gets funded. The first tells you if this is real. The second and third tell you whether $12.9 billion bought a moat or a melting asset.

    Sources

  • OpenAI Jalapeño Chip Beats Blackwell 1.9x Per Watt — Ships 2027

    The OpenAI Jalapeño chip, the company’s first custom inference ASIC, delivered 1.5x to 1.9x more AI work per watt than Nvidia’s Blackwell systems in SemiAnalysis InferenceX tests published August 25, 2026. It draws 700W against GB300’s 1,400W and cut end-to-end latency by up to 3.6x. The catch: these are engineering samples. Volume deployment does not arrive until 2027.

    What is the OpenAI Jalapeño chip?

    The OpenAI Jalapeño chip is a custom inference accelerator co-developed with Broadcom and fabricated on TSMC’s N3P node. It is built to serve tokens, not train models. OpenAI published its first third-party benchmarks this week, and they are better than any first-generation silicon has a right to be.

    The headline spec: 13.4 PFLOPS of MXFP4 compute at a 700W rating, paired with HBM4 running at 15.4 TB/s of bandwidth. In sustained operation the part draws under 550W, according to the benchmark data reported by ForkLog.

    Nvidia’s GB200 rack unit pulls 1,200W. GB300 pulls 1,400W. Rubin sits between 900W and 1,150W. Jalapeño is doing its work in roughly half the power envelope.

    The timeline is the real story

    OpenAI started design in mid-2024 and handed the chip to the fab in November 2025. That is nine months from first design to manufacturing handoff, and 16 months to tape-out — a schedule that normally takes a silicon team two to three years.

    OpenAI says its own models helped design the chip. That claim is unverifiable from the outside, but the calendar is not.

    “Jalapeño can serve more AI work per unit of power, while also returning responses more quickly,” said Richard Ho, OpenAI’s head of hardware, in comments reported by TechCrunch.

    How much faster is Jalapeño than Nvidia Blackwell?

    Across three open-weight models, Jalapeño roughly doubled Nvidia’s tokens per second per kilowatt while cutting latency by 43% to 72%. The gap widens as models get larger. On DeepSeek R1 670B, Jalapeño returned a first response in 1.65 seconds against GB300’s 5.99 seconds.

    Here are the SemiAnalysis InferenceX results as reported by ForkLog:

    Model Jalapeño (mixed TPS/kW) Nvidia system Nvidia (mixed TPS/kW) Jalapeño latency Nvidia latency
    GPT-OSS 120B 85,448 GB200 44,960 1.03s 1.80s
    DeepSeek R1 670B 19,641 GB300 11,781 1.65s 5.99s
    Kimi K2.5 1T 18,195 GB300 11,862 1.56s 5.31s

    On single-user throughput, Jalapeño hit roughly 1,400 tokens per second on GPT-OSS 120B and over 700 tokens per second on DeepSeek R1 670B.

    The aggregate claims are wider still: 1.7x to 3.6x lower end-to-end latency and 2.1x to 4.1x higher performance on interactive workloads, per The Decoder. At matched decoding speeds, The Decoder reported token-throughput-per-kilowatt advantages of 54x to 104x — a number that only makes sense in the narrow regime where GPU batching collapses.

    What SemiAnalysis actually said

    “Usually first generation chips aren’t competitive, but OpenAI is beating Nvidia Blackwell and even Rubin,” SemiAnalysis CEO Dylan Patel said, per The Decoder.

    That is a strong endorsement from an analyst house that sells research to the same hyperscalers buying Nvidia racks. Take it seriously. Take it with salt.

    Why does performance per watt decide who wins?

    Because power, not silicon, is the binding constraint on AI buildouts in 2026. Data center operators are queuing for grid interconnects measured in years. If a chip does the same work at half the watts, the same substation serves twice the revenue.

    That math is why custom ASICs keep appearing. Every watt saved on inference is a watt available for a paying customer, and inference is now the majority of frontier-lab compute spend.

    OpenAI CFO Sarah Friar framed it in cost terms: custom chips give the company “greater control over inference costs” and let it match hardware to specific tasks. Friar also said the chip “complements” existing partnerships rather than replacing them — corporate language for we are still buying your GPUs, please keep taking our calls.

    We covered the same power-and-memory squeeze from the supply side in our piece on the Nvidia AI server price hike, and the economics of fast inference in Cerebras vs Groq.

    What does this do to Nvidia’s margins?

    Nothing this quarter. Nvidia reported Q2 fiscal 2027 revenue of $96.22 billion on August 26, beating the $92.07 billion consensus, with data center revenue of $89.02 billion — up 117% year over year, according to 24/7 Wall St. EPS came in at $2.22 against a $2.09 estimate.

    Guidance was louder than the beat. Nvidia guided Q3 to $108 billion plus or minus 2%, with non-GAAP gross margins near 74% and no China data center compute revenue assumed.

    “AI has reached its inflection point. It’s doing useful work. Its tokens are productive and profitable. Now, compute is revenue,” CEO Jensen Huang said on the call.

    Nvidia also disclosed supply commitments of $279 billion, largely for Vera Rubin memory. That is a company buying ahead, not one bracing for demand loss.

    The threat is 2028, not 2026

    Custom silicon does not eat Nvidia’s revenue. It eats Nvidia’s pricing power. A 74% gross margin exists because there is no substitute at scale. Jalapeño is the first credible substitute built by Nvidia’s single largest customer.

    NVDA closed at $213.05 before the print, down 3.04% on the week and up 14.37% year to date, per 24/7 Wall St. The stock has fallen after four of its last five earnings reports despite beating consensus three quarters running.

    Who wins and who loses financially?

    Broadcom is the clearest winner. It gets ASIC design revenue, a marquee reference customer, and validation that its custom-silicon business can beat the merchant-GPU incumbent on a first attempt. Nvidia is the clearest medium-term loser, though the damage lands in 2028 pricing, not 2026 volume.

    • Broadcom — books high-margin custom ASIC revenue and proves the model. We covered its financing appetite in the Broadcom AI debt deal.
    • TSMC — wins either way. N3P wafers are N3P wafers, whether the logo says Nvidia or OpenAI.
    • HBM suppliers — Jalapeño uses HBM4 at 15.4 TB/s. More custom chips means more high-bandwidth memory demand, not less.
    • OpenAI — gains leverage in every future GPU negotiation, which may be worth more than the chip itself. Its Nvidia relationship already shifted once, as we noted when Nvidia cut its OpenAI data center guarantee.
    • Nvidia — keeps the volume through 2027, then defends 74% margins against a credible in-house alternative.
    • Second-tier inference clouds — squeezed hardest. They rent GPUs at market rates and cannot design their own.

    What’s the catch with the Jalapeño benchmarks?

    Three catches, and they matter. Jalapeño exists as engineering samples only. Rubin is already shipping to customers. And the benchmark set was chosen by the chip’s owner, run on three open-weight models, with two of Nvidia’s standard optimizations absent from the comparison.

    The Decoder reported that Jalapeño lacks multi-token prediction and speculative decoding optimizations. Those are exactly the techniques that close latency gaps on GPUs. Adding them later helps Jalapeño; adding them to the comparison today would narrow the gap.

    The models tested were GPT-OSS 120B, DeepSeek R1 670B and Kimi K2.5 1T. Larger current-generation models — DeepSeek V4 Pro, Kimi K3 — were not tested at all. Neither, notably, was any GPT-5-class OpenAI frontier model, which is the workload the chip actually has to serve.

    And the deployment schedule is honest about itself: very small volumes at the end of 2026, meaningful volume in 2027. OpenAI says a second generation is in advanced development and a third is in design.

    A chip that wins benchmarks in August 2026 must still win against whatever Nvidia ships in 2027. That is a different race.

    Frequently asked questions

    Is the OpenAI Jalapeño chip available to buy?

    No. It is an internal accelerator for OpenAI’s own inference fleet, currently at engineering-sample stage. Small-volume deployment starts at the end of 2026, with wider rollout in 2027. There is no external sales channel announced.

    Who manufactures the Jalapeño chip?

    Broadcom co-developed it with OpenAI, and TSMC fabricates it on the N3P process node. The benchmarked silicon is B0 stepping, meaning at least one revision past first tape-out.

    Does Jalapeño beat Nvidia’s Rubin?

    On the perf-per-watt figures SemiAnalysis published, yes — 1.5x to 1.9x. But Rubin is shipping to paying customers now and Jalapeño is not, so the comparison is between a product and a prototype.

    Can Jalapeño train models?

    No. It is an inference-only design. OpenAI still needs GPUs for training, which is why CFO Sarah Friar described the chip as complementing rather than replacing existing supplier relationships.

    How much power does Jalapeño use?

    It is rated at 700W and reportedly sustains under 550W in operation. Nvidia’s GB200 draws 1,200W and GB300 draws 1,400W, so Jalapeño operates in roughly half the envelope.

    Did Nvidia’s earnings show any damage from custom chips?

    None yet. Data center revenue grew 117% year over year to $89.02 billion and Q3 guidance is $108 billion. Custom silicon is a 2028 margin question, not a 2026 revenue question.

    What benchmark was used?

    SemiAnalysis InferenceX, which measures mixed tokens per second per kilowatt alongside end-to-end latency. It is a third-party benchmark, but the model selection and test configuration came from the chip’s owner.

    The bottom line

    Jalapeño is the most serious first-generation AI accelerator anyone has produced, and the power numbers are the part that should worry Nvidia. Half the watts for double the tokens is not a rounding error; it is a structural argument for custom silicon at every lab large enough to fund a design team.

    But the trade here is not “sell Nvidia.” Nvidia just printed $96.22 billion in a quarter and guided to $108 billion. The trade is that Nvidia’s 74% gross margin now has an expiry date attached, and the market will start pricing that date long before 2028 arrives.

    The honest read: OpenAI has proven it can build a chip. It has not yet proven it can build ten million of them, on schedule, while Nvidia iterates annually. Benchmarks are cheap. Yield is not.

    Sources

  • Nvidia AI Server Price Hike: 15% More, and Memory Is Why

    Nvidia is raising AI server prices by more than 15% on Grace Blackwell and Vera Rubin systems shipping in early 2027, Bloomberg reported on August 24, 2026. Memory is the reason. UBS puts memory at 62% of a Vera Rubin superchip’s $38,902 bill of materials, up from 53% on Grace Blackwell. TrendForce estimates the hike adds at least $5 billion to a 1-gigawatt data center.

    For three years the AI trade had one simple rule: Nvidia sets the price, and everyone pays it. That rule still holds. What changed is who Nvidia is paying.

    The Nvidia AI server price hike is not a margin grab. It is a pass-through. And the numbers underneath it say the memory makers, not the GPU designer, now control the cost curve of the AI build-out.

    How much is Nvidia raising AI server prices?

    More than 15% in many cases, effective on systems shipped early next year. Bloomberg reported the increases on August 24, citing people familiar with the matter. TrendForce, summarizing the same reporting, said some configurations could reach 17%. Nvidia did not respond to requests for comment.

    The warnings did not go to the cloud giants directly. According to Bloomberg, Nvidia notified the contract server manufacturers that assemble systems for Microsoft, Alphabet’s Google, and Oracle.

    That routing matters. The ODMs absorb the notice first, then reprice their own quotes. The cloud buyers find out when the invoice changes.

    Which systems are affected

    • Grace Blackwell systems — the current generation, still shipping in volume.
    • Vera Rubin systems — the next generation, with first shipments in early 2027.
    • Increases vary by chip generation and by memory configuration, per Bloomberg. Denser memory builds take the larger hit.

    Why is Nvidia raising prices now?

    Because memory has gone from a line item to the line item. Morgan Stanley estimates GPU silicon has fallen from more than 80% of AI server cost to roughly half that level in next-generation systems. The gap did not close because GPUs got cheaper. It closed because DRAM got expensive.

    A Vera Rubin NVL72 rack carries 74.7 TB of DRAM — 20.7 TB of HBM4 plus 54 TB of LPDDR5X, according to UBS’s teardown. That is the DRAM content of roughly 4,500 smartphones in a single rack.

    Every one of those bits is bought in the tightest memory market in a decade.

    What UBS found inside a Vera Rubin superchip

    UBS’s bill-of-materials analysis is the clearest picture available of where the money actually goes.

    Component Cost Share of superchip
    Total Vera Rubin superchip $38,902 100%
    All memory $24,297 62%
    SOCAMM2 LPDDR5X $19,355 49.8%
    HBM4 $4,943 12.7%
    Everything else $14,605 38%

    On Grace Blackwell, UBS put memory at 53% of cost. On Vera Rubin it is 62%, and the absolute memory bill rose about 2.5x between generations.

    One caveat worth holding onto: that 2.5x blends two different things. Vera Rubin carries more memory and pays more per gigabyte. It is not a pure price signal.

    How much has DRAM actually gone up?

    Steeply, and for longer than most forecasts allowed. TrendForce data cited by Tom’s Hardware shows conventional DRAM contract prices rising 90–95% quarter-over-quarter in Q1 2026 and a projected 58–63% in Q2 2026. Server DRAM is expected to climb every quarter through the second half of 2027.

    The consumer market tells the same story in plainer numbers. A mainstream 32GB DDR5-6000 kit runs about $392 today against $110–$140 a year ago, per Tom’s Hardware.

    Supply was committed early. SK hynix had sold out its entire 2026 production capacity by October 2025. Samsung and SK hynix raised 2026 HBM3E prices by roughly 20%.

    And HBM makes the squeeze worse mechanically: it consumes roughly four times the wafer area of conventional DRAM per bit shipped. Every HBM4 order crowds out ordinary server memory on the same fab.

    Who profits from the Nvidia AI server price hike?

    Not Nvidia, on the arithmetic. The memory suppliers capture the increase, the ODMs pass it through, and the hyperscalers eat it. Nvidia’s role here is closer to toll collector than beneficiary — and its own gross margin may be the quiet casualty.

    Work the math. Nvidia runs roughly a 75% gross margin, so the bill of materials is about 25% of the sale price. If memory is 62% of that BOM, memory is about 15.5% of the price. A 2.5x memory cost increase adds roughly 23 points of price to cost.

    A 15% price hike does not cover 23 points. Something has to give.

    Three readings, and the market has not settled on one:

    1. The 15% is an opening installment. More increases follow as 2027 contracts reprice.
    2. Nvidia is absorbing the difference. Gross margin drifts from ~75% toward the high 60s.
    3. The 2.5x is generational, not inflationary. Higher memory content is sold at a higher system ASP, so the comparison overstates the pass-through problem.

    Reading three is the most likely and the least discussed. It is also the one that would let Nvidia keep its margin story intact — which is precisely why it deserves scrutiny rather than acceptance.

    Why this matters for AI capex

    Because it reprices the entire build-out. TrendForce estimates the increase adds at least $5 billion to the cost of a 1-gigawatt AI data center. At the scale hyperscalers are now committing to, that is not a rounding error — it is a line in the capital plan that did not exist last quarter.

    The second-order effects are where this gets interesting.

    The uncomfortable version: AI compute has been getting cheaper per unit of intelligence for three straight years. This is the first credible input cost that pushes the other way.

    This post is reporting and analysis, not financial advice.

    Frequently asked questions

    How much is Nvidia raising AI server prices?

    More than 15% in many cases, with some configurations reaching 17% per TrendForce. Increases vary by chip generation and memory configuration.

    When do the new prices take effect?

    On systems shipped in early 2027, according to Bloomberg’s August 24, 2026 report.

    Which Nvidia systems are affected?

    Grace Blackwell and Vera Rubin server systems. Both are rack-scale platforms sold to cloud and enterprise data center operators.

    Why are AI server prices going up?

    Memory costs. UBS puts memory at 62% of a Vera Rubin superchip’s cost, and DRAM contract prices have risen every quarter through 2026 amid an HBM-driven supply squeeze.

    Who was notified about the price increases?

    Contract server manufacturers that build systems for Microsoft, Google, and Oracle, per Bloomberg. Nvidia did not comment publicly.

    How much does this add to a data center?

    TrendForce estimates at least $5 billion in additional cost for a 1-gigawatt AI data center.

    Does this hurt Nvidia’s margins?

    Possibly. Nvidia runs roughly a 75% gross margin. If memory costs rose 2.5x generationally, a 15% price increase may not fully offset it — though part of that increase reflects more memory content per system, not pure inflation.

    The bottom line

    The Nvidia AI server price hike is the clearest sign yet that the AI supply chain’s power center is shifting. For three years the scarce input was GPU wafer allocation. In 2027 it is memory, and the companies that own it — SK hynix, Samsung, Micron — are the ones setting terms.

    Watch two things next. First, whether Nvidia’s gross margin guidance holds through the fiscal year, because that is where the pass-through gap shows up. Second, whether any hyperscaler publicly revises a gigawatt commitment. The first cost-driven downgrade of an announced buildout would tell you the memory squeeze has stopped being an engineering problem and started being a financial one.

    Sources

  • Etched Valuation Doubles to $21 Billion in Under a Month

    Etched raised $700 million at a $21 billion post-money valuation on August 18, 2026, led by quant trading firm Jane Street — which is also its first paying customer. The Etched valuation doubled from $10.3 billion in July, according to TechCrunch. The transformer-ASIC startup has booked more than $1 billion in signed orders and has now raised close to $2 billion in total.

    A four-year-old chip company just repriced itself faster than almost anything in the AI hardware cycle. The question is whether the order book justifies it.

    How much did Etched raise, and at what valuation?

    Etched raised $700 million at a $21 billion post-money valuation, announced Tuesday, August 18, 2026. Jane Street led the round. That is a doubling from the $10.3 billion Series C the company closed in July 2026 — roughly four weeks earlier, per TechCrunch’s reporting.

    The step-up is the headline. In December 2025 Etched was worth $5 billion. Eight months later it is worth $21 billion, a 4.2x move without a single public revenue disclosure.

    The Etched valuation timeline

    Date Round Valuation Lead investor
    December 2025 Series B extension $5 billion Not disclosed
    July 2026 Series C, $300M $10.3 billion Sequoia Capital
    August 18, 2026 $700M round $21 billion Jane Street

    Total capital raised is now close to $2 billion, according to Tech Startups. The cap table includes Sequoia Capital, Kleiner Perkins, Andreessen Horowitz, Peter Thiel, Tiger Global, Bain Capital Ventures, Stripes, Primary, Positive Sum and Blackstone.

    Why is Jane Street both the lead investor and the first customer?

    Jane Street tested Etched’s system, installed a rack in its own datacenter, and then led the round. That dual role is the most important detail in the announcement — and the one that deserves the most scrutiny. A lead investor who is also the reference customer validates the product and inflates the comparable at the same time.

    “Etched’s unique approach to inference delivers the precision we will need to support our most demanding workloads,” Jane Street said in the announcement, adding that it has “our own rack running in our datacenter.”

    Quant trading is an unusually favorable first market. Latency is worth real money there, the workloads are narrow and stable, and the buyer has no procurement committee. Whether that translates to hyperscalers running heterogeneous model fleets is a genuinely open question.

    What Etched actually ships

    The company sells what it calls frontier inference clusters, built around two custom components: a low-voltage prefill chip and cluster-scale memory sized for the decode phase. Its Sohu part is marketed as the world’s first transformer ASIC.

    Etched says it went from receiving test silicon at TSMC to running inference workloads in 44 days — fast for a first-silicon bring-up, where months is normal.

    What is a transformer ASIC, and why does it threaten Nvidia?

    A transformer ASIC hard-codes one model architecture into silicon instead of staying programmable. You lose flexibility and gain throughput and power efficiency. The bet is that transformers stay dominant long enough for fixed-function chips to pay back their tape-out cost before the architecture moves.

    Nvidia’s moat is generality plus CUDA. An ASIC attacks exactly the workload where generality is least valuable: high-volume, steady-state inference of one model family.

    Etched has also been buying the expertise directly. Roughly 15% of its ~400 employees came from Nvidia — about 60 people, per Tech Startups. Systems engineer Brian Loiler, who spent 23 years at Nvidia before joining in 2024, recruited around a dozen more Nvidia engineers; some turned down counteroffers.

    Founded in 2022 by Harvard dropouts Gavin Uberti and Chris Zhu, the company operates from San Jose with an internal datacenter. Sequoia GP Sonya Huang summed up the historical skepticism the round is arguing against: “Don’t back the kids in chips.”

    Who else is buying into custom AI silicon right now?

    Etched is not an outlier — it is the loudest datapoint in a two-week run of custom-silicon deals. Three of the largest chip buyers and builders in the market all moved on inference-specific hardware in August 2026.

    • AMD acquired Taalas on August 6, 2026, for an undisclosed sum. Taalas etches models directly into silicon; AMD says it will fold the technology into its accelerator roadmap alongside Instinct GPUs.
    • Marvell granted Google a warrant for up to 58.97 million shares — worth up to $12.2 billion — tied largely to purchasing targets through fiscal 2033. Marvell stock jumped more than 11% in premarket trading on August 19.
    • Nvidia itself is hedging, spending on the layer below the chip: it took a stake in datacenter developer Cloverleaf and, as we covered, paid roughly $6 billion for Poolside’s model factory and 109 staff.

    The pattern is consistent. Everyone with capital is buying inference efficiency, and they are paying acquisition-grade prices for it. The same dynamic is visible in the fast-inference API market, where specialized silicon is already competing on price per token.

    Why this matters for the AI market and investors

    Inference, not training, is now where the compute money goes — and that is the market Etched is built for. Nvidia posted $81.6 billion in total revenue in Q1 FY2027, with $75.2 billion from Data Center, up 92% year over year, according to the company’s own results release for the quarter ended April 26, 2026.

    Against that, $21 billion for a startup with roughly $1 billion in signed orders is a bet on share shift, not on displacement. Etched would need to compound for years to matter to Nvidia’s income statement.

    The more useful read is directional. Capital is repricing the assumption that general-purpose GPUs capture all inference margin. That assumption also underwrites the debt now funding the buildout — see the Broadcom financing package and Nvidia’s decision to cut its OpenAI datacenter guarantee from $250 billion to $120 billion.

    The skeptical case

    A valuation that doubles in four weeks on the same order book is a financing event, not an operating one. Nothing in the disclosed figures changed between July and August except who was writing the check.

    Three specific risks:

    1. Architecture risk. A transformer ASIC is a leveraged bet that transformers stay dominant. If the frontier moves to a materially different architecture, the silicon does not follow.
    2. Customer concentration. One named customer, who is also the lead investor. Signed orders above $1 billion are unaudited and self-reported.
    3. Incumbent response. Nvidia has $75.2 billion of quarterly Data Center revenue to defend a niche with, and AMD just bought a direct competitor to Etched’s approach.

    This post is reporting and analysis, not financial advice.

    Frequently asked questions

    What is the Etched valuation now?

    $21 billion post-money, set by the $700 million round announced August 18, 2026.

    Who led Etched’s $700 million round?

    Jane Street, the quantitative trading firm, which is also Etched’s first paying customer and has a rack running in its own datacenter.

    How much has Etched raised in total?

    Close to $2 billion, according to Tech Startups. The prior round was a $300 million Series C at $10.3 billion in July 2026.

    What does Etched’s Sohu chip do?

    Sohu is marketed as the world’s first transformer ASIC — silicon purpose-built for transformer inference rather than general-purpose computation, trading flexibility for throughput and power efficiency.

    How many Etched employees came from Nvidia?

    About 15% of roughly 400 employees, or around 60 people, per Tech Startups.

    Is Etched profitable?

    The company has not disclosed revenue or profitability. It reports more than $1 billion in signed orders and has begun shipping chips.

    Who are Etched’s founders?

    Gavin Uberti and Chris Zhu, Harvard dropouts who founded the company in 2022. Robert Wachen is co-founder and COO.

    The bottom line

    Etched has the most aggressive valuation trajectory in AI hardware and a real product shipping into a real datacenter. It also has one named customer who set the price.

    Watch two things next: whether a hyperscaler or a frontier lab signs, and whether the $1 billion order book converts to disclosed revenue. If a second, unaffiliated buyer of scale appears in the next two quarters, $21 billion looks early. If it does not, this round will read as the moment ASIC enthusiasm outran ASIC demand.

    Sources

  • Nvidia AVO Hits 100% on ARC-AGI-3. The Model Alone Scored 30.2%.

    Nvidia AVO — Agentic Variation Operators — scored 100.00 RHAE on the ARC-AGI-3 public set on August 21, 2026, clearing all 183 levels across 25 environments in 6,624 actions, roughly 12% fewer than the VISTA baseline. The same base model, Claude Opus 5, scores 30.2% on its own. The harness did the work, and that changes where agent money goes.

    What is Nvidia AVO?

    Nvidia AVO stands for Agentic Variation Operators. It is not a model. It is a general-purpose coding-agent system that wraps an existing frontier model in a loop — inspect, plan, implement, evaluate — plus persistent memory and a supervisor that intervenes when progress stalls. Nvidia published the results on August 21, 2026.

    The base model inside the winning run was Anthropic’s Claude Opus 5. Nvidia also ran limited experiments with GPT-5.6 Sol on a subset of games, and labeled those findings preliminary.

    That detail is the whole story. Nvidia did not train a better reasoner. It built better scaffolding around someone else’s reasoner.

    How the AVO loop works

    AVO runs a four-step cycle: inspect the current context, plan a change, implement it, then evaluate the result against the environment.

    Two additions separate it from a standard agent loop, according to Nvidia’s technical blog:

    • Persistent memory that carries forward prior implementations, evaluation results and reasoning across the whole run, not just the current context window.
    • A supervision mechanism that watches the trajectory and redirects the agent when it detects the run has stopped making progress.
    • Variation operators that generate structured alternatives rather than retrying the same failed approach.

    Why the supervisor is the expensive part

    Long-horizon agent failure is rarely a single wrong answer. It is a slow drift — the agent loops on a dead approach and burns tokens without noticing.

    A supervisor that detects stagnation is cheap to describe and hard to build. It is also the component least likely to transfer cleanly to another benchmark.

    What is ARC-AGI-3 and why does a 100% score matter?

    ARC-AGI-3 is ARC Prize’s interactive reasoning benchmark: 25 pixel-art puzzle environments containing 183 public levels. Agents get no instructions, no rules and no goal labels. They must infer the mechanics purely by playing. When the benchmark launched, humans cleared 100% of environments and the best AI managed 0.37%.

    That 0.37% figure is why this result registered. ARC-AGI-3 was designed as the benchmark models could not touch.

    How RHAE scoring works

    The metric is RHAE — Relative Human Action Efficiency. It combines task completion with how many actions the agent needed per level, measured against initial human performance, then aggregates across every level and environment.

    So a 100.00 does not just mean “finished everything.” It means finishing everything at roughly human action efficiency. ARC Prize published its human performance dataset specifically so this number would have a floor to sit on.

    How much did the harness add versus the raw model?

    The gap is 30.2% to 100.00 — the same model class, wrapped differently. ARC Prize reported Claude Opus 5 at 30.2% on ARC-AGI-3 in July 2026, which it called a genuine reasoning leap at the time. Nvidia’s harness took that model to a clean sweep of the public set.

    SystemARC-AGI-3 resultActions usedReported by
    Best AI at benchmark launch0.37%ARC Prize
    Claude Opus 5 (bare model)30.2%ARC Prize, July 2026
    VISTA baseline agentCleared same level sets7,542Nvidia
    Nvidia AVO (Claude Opus 5 inside)100.00 RHAE, all 183 levels6,624Nvidia, Aug 21 2026

    For context on the base model’s ceiling elsewhere: Claude Opus 5 at maximum reasoning effort scores 97.5% on ARC-AGI-1 and 90.4% on ARC-AGI-2, per The New Stack. ARC-AGI-3 was the one that stayed hard.

    Nvidia’s own framing, from the blog post: “The model matters, but the model is not the entire agent.”

    Is the 100% score real, or is it benchmark theater?

    It is real, and it is narrower than the headline suggests. The score covers the ARC-AGI-3 public set only — not the semi-private or private competition sets that ARC Prize maintains precisely to catch overfitting. Nvidia says so in its own post.

    The public set is not the hidden exam

    Benchmark designers split datasets for a reason. A public set is a practice test with the answers eventually leaking into the ecosystem around it.

    One commenter on Nvidia’s announcement put it bluntly, as flagged in explainX’s write-up: “I would not file this as solved AGI… if you post 100 like it is the hidden exam.” Until AVO posts a semi-private number, that objection stands.

    These are not controlled ablations

    Nvidia explicitly labels its comparisons as not controlled ablations. That matters more than it sounds.

    The AVO-versus-VISTA action count — 6,624 against 7,542 — varies agent backends, observation formats, memory systems and reasoning settings all at once. The 100.00-versus-30.2% comparison swaps the entire system architecture and the reasoning-effort setting simultaneously.

    Neither number isolates how much the harness itself contributed. The honest reading is “a well-built harness closed a very large gap,” not “the harness is worth exactly 70 points.”

    Nvidia also disclosed no compute cost, no token usage and no wall-clock runtime for the 6,624 actions. For anyone pricing an agent product, that is the number that actually matters — and it is missing.

    Has AVO done anything useful outside a puzzle benchmark?

    Yes, and this is the part investors should read twice. Nvidia ran AVO continuously for seven days on GPU-kernel optimization, exploring more than 500 optimization directions. The system produced kernels that beat FlashAttention-4 by up to 10.5% on NVIDIA DGX B200 hardware.

    FlashAttention is not a soft target. It is hand-tuned infrastructure that the entire industry’s inference economics rest on.

    A 10.5% kernel improvement compounds across every token served on that hardware. If it holds in production, it is worth more to Nvidia than the benchmark headline — and it lands in the same week the company has been buying capability outright elsewhere.

    Who wins and who loses financially?

    The winner is whoever owns the orchestration layer. If a 30% model becomes a 100% agent through harness design, then value is accruing above the weights, not inside them. That is bad news for anyone whose entire moat is a checkpoint.

    Nvidia is climbing the stack

    AVO did not appear in isolation. On the same day, Nvidia paid $6 billion to license Poolside’s model-development software and invested $1 billion more in the startup, according to PYMNTS.

    A chip company publishing frontier agent architecture and licensing a model factory in the same 24 hours is not a coincidence. It is a company that has watched its customers capture the margin its silicon creates — the same dynamic behind its recalculated OpenAI data center guarantee.

    Model labs keep pricing power, for now

    Note who supplied the brain: Anthropic. AVO’s best run needed Claude Opus 5, and the harness could not manufacture reasoning that was not already there — the 0.37% launch-day figure is proof that scaffolding alone does nothing on a weak model.

    So frontier labs still sell the scarce input. What they lose is the claim that the model is the product, which shows up quickly in cheaper models closing capability gaps.

    Agent startups just got a harder question

    Three practical implications for anyone building or funding an agent company:

    1. Harness design has not hit diminishing returns. A 30-to-100 jump says the scaffolding layer is still under-engineered — which is opportunity and commoditization risk in the same sentence.
    2. Your differentiator may be a blog post away from replication. Persistent memory plus a stagnation supervisor is a describable architecture, not a trade secret.
    3. Nvidia is now a potential competitor, not just a supplier. It has the hardware, the capital, and as of August 21, published frontier agent research.

    The cost question decides all three. Running a supervised, memory-heavy loop for 6,624 actions is not free, and the economics look very different depending on whether the underlying tokens cost $2 or $60 per million — the same math that drives coding-agent unit costs and inference vendor selection.

    Frequently asked questions about Nvidia AVO

    Is Nvidia AVO a new AI model?

    No. AVO is an agent system — a harness — that runs on top of existing frontier models. The reported 100.00 RHAE run used Claude Opus 5 as its base model.

    Did Nvidia AVO solve AGI?

    No. The score covers ARC-AGI-3’s public set of 183 levels across 25 environments. ARC Prize also maintains semi-private and private sets, and AVO has not posted a result on those.

    What does RHAE mean?

    Relative Human Action Efficiency. It scores both whether an agent completes a level and how many actions it needed relative to initial human performance, aggregated across the benchmark.

    Can developers use AVO today?

    Nvidia’s August 21 post describes the architecture and results. It does not announce a code or weights release, so treat AVO as published research rather than a shippable dependency.

    How much does an AVO run cost?

    Nvidia did not disclose compute cost, token usage or wall-clock time for the benchmark run. Without those figures, the result cannot be compared on a cost-per-task basis against cheaper agent harnesses.

    What was the FlashAttention-4 result?

    Running for seven days across 500-plus optimization directions, AVO produced GPU kernels that outperformed FlashAttention-4 by up to 10.5% on NVIDIA DGX B200 hardware.

    Does this make Claude Opus 5 look better or worse?

    Both. The model was capable enough to be driven to 100.00 RHAE, and weak enough on its own to score 30.2%. The delta belongs to the harness, not the checkpoint.

    The bottom line

    Nvidia AVO is the most important agent result of the month, and the headline number is the least interesting part of it.

    A 100.00 on a public set with no controlled ablations and no disclosed cost is a demonstration, not a benchmark victory. Anyone treating it as “ARC-AGI-3 is solved” is reading a press release as a result.

    What survives scrutiny is the gap: 30.2% to 100.00, same model, different scaffolding. That gap is the clearest evidence yet that in 2026 the agent layer, not the model layer, is where the remaining engineering leverage sits.

    And the FlashAttention-4 kernels are the tell. Nvidia did not build AVO to win a puzzle leaderboard. It built AVO to make its own hardware faster — and, at a moment when record sums are being raised to finance AI chip capacity, to stop being only the company that sells the machines.

    Sources

  • Nvidia Poolside Deal: $6 Billion for a Model Factory and 109 Staff

    Nvidia is paying Poolside $6 billion to license its model-building software and hiring 109 of the startup’s staff, according to Newcomer, which broke the story on August 20, 2026. A separate $1 billion investment values what remains at $12 billion pre-money. Nvidia shares closed the week down roughly 5%. No company legally changes hands.

    The Nvidia Poolside deal is the third time in twelve months that the world’s most valuable chipmaker has bought a startup without buying a startup. It is becoming a template.

    What exactly is the Nvidia Poolside deal?

    Three transactions in one package. Nvidia pays $6 billion for a non-exclusive license to Poolside’s “model factory,” extends offers to 109 employees, and invests $1 billion at a $12 billion pre-money valuation. Poolside keeps its name, its three founders, and its corporate independence.

    The “model factory” is not a model. It is the system Poolside built to produce models — the training pipeline, the data infrastructure, the orchestration layer.

    Nvidia is buying the assembly line, not the car.

    Bloomberg confirmed the terms on August 20, citing Newcomer’s reporting. The Information reported the same package the following day.

    The deal terms, line by line

    Component Terms Source
    Technology license $6 billion, non-exclusive Newcomer, Aug 20, 2026
    Equity investment $1 billion Bloomberg
    Valuation of remaining entity $12 billion pre-money Newcomer
    Staff receiving Nvidia offers 109 employees The Next Web
    Founders staying with Poolside 3 The Next Web
    Proceeds distributed to investors By end of 2027 Poolside investor letter
    Prior Nvidia commitment to Poolside Up to $1 billion (October 2025) The Next Web

    Why did Poolside sell its model factory?

    Because it could not afford the chips. Poolside’s investor letter, quoted by The Next Web, describes a financing failure with a hard deadline: the company needed $2 billion in six weeks to pay for a 40,000-GPU cluster, missed the window, and lost the allocation.

    The letter is unusually blunt. “We had a 6 week window in which to raise $2 billion dollars to pay for a 40,000 GB300 cluster coming online in January,” it reads. “We didn’t close it in time, and we lost the cluster.”

    Poolside’s own assessment: a frontier-competitive model needs 10,000 to 20,000 of those chips today, and materially more next year.

    That is the whole story of the independent model lab in 2026, compressed into two sentences. The research talent is not the constraint. The capital stack is.

    • The gap: $2 billion needed in six weeks, against a $12 billion pre-money valuation
    • The consequence: allocation forfeited, frontier ambitions shelved
    • The pivot: Poolside moved from coding agents into data center operations and open-weight releases before the deal
    • The buyer: the company that sells the chips it could not pay for

    CEO Eiso Kant and two co-founders remain. The people who actually built the thing — fewer than 70 on the model itself, under 115 across engineering and research combined — largely go to Nvidia.

    How does this compare to Nvidia’s Groq and Enfabrica deals?

    It is the same structure at a different price. Across three transactions, Nvidia has committed roughly $27 billion to license technology and absorb teams while leaving the original corporate entities standing. Groq was the largest at about $20 billion. Enfabrica was roughly $900 million.

    Target Reported value What Nvidia received
    Enfabrica ~$900 million License plus networking team
    Groq ~$20 billion Non-exclusive design license, founders, most staff
    Poolside $6B license + $1B equity Model-factory license, 109 staff
    Combined ~$27 billion Three teams, zero acquisitions

    The pattern is deliberate enough that it now has a name in the trade press: the reverse acquihire. Buy the license, hire the people, leave the shell.

    Is the reverse acquihire an antitrust workaround?

    Two US senators have already said so in writing. On March 23, 2026, Elizabeth Warren and Richard Blumenthal wrote to Jensen Huang about the Groq deal, arguing that Nvidia “has effectively acquired Groq in all but name” by licensing its technology and hiring its key employees.

    The letter cites Nvidia’s roughly 90% share of the GPU market and warns the structure “could stifle competition, further entrenching NVIDIA’s dominance in the AI chip industry.”

    The senators’ core objection is procedural. A conventional acquisition triggers premerger notification and agency review. A license plus a hiring spree does not — even when the economic result is indistinguishable.

    The FTC and DOJ retain authority to investigate consummated transactions regardless of filing status. Whether they will is a different question. As of this writing, no public enforcement action has been announced against any of the three deals.

    Read the full Warren-Blumenthal letter for the argument in the senators’ own words.

    What are the skeptical questions about the $6 billion price?

    Start with the arithmetic. Nvidia is paying $6 billion for a non-exclusive license to software built by fewer than 115 people at a company that just failed to raise $2 billion. That is roughly $55 million per engineer hired, and the license does not stop Poolside from licensing the same technology elsewhere.

    Non-exclusive is the word doing the most work in this deal.

    Second question: what is Nvidia actually short of? It is not model-training expertise — Nvidia has built plenty. The likelier answer is speed. Buying a working pipeline compresses years into a quarter.

    Third: the money moves in a familiar circle. Nvidia committed up to $1 billion to Poolside in October 2025. Poolside spent on Nvidia hardware. Nvidia now pays $6 billion back, some of which flows to investors by end of 2027, and takes another $1 billion equity position. Revenue and investment are increasingly hard to separate on this balance sheet.

    We flagged the same circularity concern when Nvidia cut its OpenAI data center guarantee from $250 billion to $120 billion — a revision that suggested even Nvidia has limits on how much demand it will underwrite itself.

    The market noticed. Nvidia shares fell about 5% over the week of the announcement, closing Friday down 0.9%, though the stock remains up 14.5% year to date.

    Why this matters for the AI market

    The Nvidia Poolside deal marks the point where compute access stopped being a competitive advantage and became a gate. Poolside had the talent, the models, and a $12 billion valuation. It still could not clear a $2 billion payment on schedule, and that alone ended its frontier ambitions.

    For investors, three implications follow.

    1. The exit landscape has changed. A reverse acquihire returns capital without an acquisition premium, an IPO, or regulatory review. Cap tables should price that in.
    2. Valuation and viability have decoupled. A $12 billion paper valuation did not translate into $2 billion of callable cash in six weeks.
    3. Nvidia is consolidating the stack quietly. Roughly $27 billion across three deals, none of which required a merger filing.

    Compare this to the conventional route: Stripe paid an estimated $7 billion and actually bought the company when it acquired OpenRouter. Nvidia is getting comparable strategic value for less, with less scrutiny.

    Meanwhile the debt markets are doing their own version of the same trade — Broadcom is arranging up to $100 billion to finance AI chip infrastructure. Capital is chasing compute from every direction at once.

    This post is reporting and analysis, not financial advice.

    Frequently asked questions about the Nvidia Poolside deal

    Short answers to the questions readers are asking about the structure, the price, and what happens next.

    Did Nvidia acquire Poolside?

    No. Poolside remains an independent company with its three founders and its own board. Nvidia licensed technology and hired staff. No change of control occurred.

    How much is Nvidia paying in total?

    $6 billion for the non-exclusive license plus a $1 billion equity investment — $7 billion combined, per Newcomer and Bloomberg reporting from August 20, 2026.

    How many Poolside employees are joining Nvidia?

    109 received offers, according to The Next Web. Poolside had fewer than 115 people across engineering and research in total.

    What is a “model factory”?

    The infrastructure and process Poolside built to train AI models — pipelines, data systems, orchestration. Nvidia licensed the production system rather than any individual model.

    Why is this called a reverse acquihire?

    A normal acquihire buys a company to get its people. Here Nvidia gets the people and the technology while the company survives, avoiding merger review.

    Has this structure faced regulatory pushback?

    Senators Warren and Blumenthal challenged Nvidia’s similar $20 billion Groq deal in a March 23, 2026 letter, calling it an acquisition “in all but name.” No enforcement action has followed publicly.

    What happens to Poolside now?

    It continues with $1 billion in fresh capital at a $12 billion pre-money valuation and plans to distribute the $6 billion license proceeds to investors by end of 2027.

    The bottom line

    Nvidia has found a way to buy companies that does not look like buying companies, and it has now used it three times for roughly $27 billion. The Poolside deal is the cleanest example yet: a startup that could not fund its own chips sold the machine that would have used them, to the company that makes them.

    Expect two things next. More labs will take this exit — the economics of independent frontier training are brutal, and a license-plus-hire returns capital fast. And expect the structure to draw a formal response from Washington, because three deals is a pattern, not a coincidence.

    Watch for whether the FTC opens a review. That is the variable that decides whether this template survives 2027.

    Sources

  • Cerebras vs Groq: Which Fast Inference API Is Worth the Money

    Cerebras vs Groq comes down to one trade. Cerebras serves GPT-OSS-120B at 1,641 tokens per second for $0.75 per million output tokens. Groq serves the same model at roughly 500 for $0.60. You pay about 25% more on output for roughly three times the speed. Buy Cerebras when a human or an agent is waiting. Buy Groq for batch work and overnight jobs.

    That trade just got sharper. On August 18, 2026, Cerebras announced the CS-4, a system it claims runs GPT-OSS-120B at more than 4,400 tokens per second per user.

    If that number survives contact with production traffic, the speed gap stops being a nice-to-have and starts being a product feature you can charge for.

    What changed in the Cerebras vs Groq race this week?

    Cerebras shipped a new generation of silicon and Nvidia now owns its main rival’s technology. Those two facts reshape the fast-inference market. The CS-4 raises Cerebras’ ceiling; the Nvidia-Groq deal means Groq’s aggressive pricing is no longer set by a scrappy independent.

    The CS-4 numbers that matter

    Per the Cerebras announcement, the CS-4 delivers 750 PFLOPS of AI compute against the CS-3’s 125 PFLOPS. Memory bandwidth jumps from 21.6 to 129.6 petabytes per second.

    The WSE-3 Turbo processor behind it packs 4 trillion transistors and 900,000 AI cores across 46,225 square millimeters of silicon, with 44GB of on-chip SRAM.

    Wafer-to-wafer latency drops from 5 microseconds to 2. Cerebras also claims up to 10x the throughput per watt versus the CS-3, and support for models above 50 trillion parameters.

    The CS-4 product page adds a second claim worth watching: more than 1,000 tokens per second on models exceeding 10 trillion parameters. First shipments began in Q3 2026.

    CTO Sean Lie framed the pitch in agent terms: “Being 30 times faster gives an agentic system room for significantly more reasoning, verification, or tool use in the same wall-clock time.”

    Why Nvidia now sits on both sides

    Groq is no longer an independent challenger. CNBC reported on December 24, 2025 that Nvidia agreed to buy Groq’s assets for about $20 billion — its largest deal on record.

    The Groq API still runs and still undercuts Cerebras. But the pricing that made Groq attractive is now a line item inside the company that also sells the GPUs Groq was built to beat.

    That matters for anyone building a business on a specific cost per token. Cerebras is the last large pure-play fast-inference vendor with its own silicon and its own incentive to keep prices down.

    How much does fast AI inference cost per million tokens?

    Cerebras is the most expensive way to run GPT-OSS-120B among mainstream providers. Groq sits mid-pack. Commodity GPU serverless tiers cost a fraction of both. The spread on the identical open-weights model is roughly 7x on output tokens, which is far wider than most teams assume.

    Here is the pricing snapshot for GPT-OSS-120B as tracked by PricePerToken on August 21, 2026, with measured speed where it is published.

    Provider Input $/M Output $/M 1M in + 1M out Measured output speed
    Cerebras $0.35 $0.75 $1.10 1,641 tok/s
    Groq $0.15 $0.60 $0.75 ~500 tok/s
    SambaNova $0.14 $0.95 $1.09 Not published
    Together AI $0.15 $0.60 $0.75 Not published
    Amazon Bedrock $0.15 $0.60 $0.75 Not published
    Baseten $0.10 $0.50 $0.60 Not published
    Google $0.09 $0.36 $0.45 Not published
    Fireworks $0.10 $0.10 $0.20 Not published
    DeepInfra $0.037 $0.170 $0.207 Not published
    OpenRouter $0.030 $0.170 $0.200 Not published
    Cerebras speed from Artificial Analysis; Groq speed from CloudZero, May 2026.

    Run a balanced million-in, million-out workload and Cerebras costs $1.10 against Groq’s $0.75 — a 47% premium. On output tokens alone the gap narrows to 25%.

    Both are expensive next to Fireworks at $0.20 or the OpenRouter route at $0.20 for the same weights.

    The discount lever most teams forget

    Groq’s list price is not its real price. CloudZero notes that Groq’s Batch API and prompt caching each cut rates by 50%, and the two stack to roughly 25% of on-demand pricing.

    Applied to GPT-OSS-120B, that pushes effective output cost toward $0.15 per million. Cerebras publishes no equivalent public stacking discount; its pricing page lists a $5 free trial, a $10 self-serve developer tier, and custom enterprise rates.

    For any workload that tolerates a batch window, Groq is not 20% cheaper. It is closer to 5x cheaper.

    Which is faster in practice, Cerebras or Groq?

    Cerebras wins today, and not narrowly. Independent measurement puts Cerebras at 1,641 tokens per second on GPT-OSS-120B with 0.46 seconds to first token. Groq’s published figure for the same model is around 500 tokens per second. That is a 3.3x throughput advantage before the CS-4 ships at volume.

    Artificial Analysis also clocks Cerebras at 1,402 tokens per second on Gemma 4 31B in reasoning mode, at a blended $0.24 per million.

    What the speed gap costs in real money

    Take a 100,000-token generation — a long agent trace or a full document rewrite.

    • Cerebras today: 61 seconds at 1,641 tok/s, costing $0.075 in output tokens.
    • Groq today: 200 seconds at 500 tok/s, costing $0.060.
    • Cerebras CS-4 claim: 23 seconds at 4,400 tok/s.
    • The math: 1.5 cents buys back 139 seconds — roughly 11 cents per minute of latency removed.

    Eleven cents a minute is trivial if a customer is watching a cursor blink. It is indefensible if the job runs at 3 a.m. and nobody reads the output until morning.

    One honest caveat: the 4,400 tok/s figure is a vendor claim tied to hardware that only started shipping this quarter. The 1,641 figure is measured on the live API. Do not budget against the former.

    Is Cerebras worth the premium for AI agents?

    Yes, for interactive and multi-step agent work — and this is the only case where the premium clearly pays. Agent loops multiply latency: ten sequential tool calls at 200 seconds each is a 33-minute task. The same loop at Cerebras speed finishes in about 10 minutes.

    That compounding is the whole argument. A single completion at 500 tokens per second feels fine. Twenty of them chained behind a task does not.

    Who should buy what

    Use case Pick Why
    Live chat, copilots, voice Cerebras 0.46s TTFT and 1,641 tok/s; latency is the product
    Multi-step agents with tool calls Cerebras Per-step latency compounds across the loop
    Overnight batch, evals, data labeling Groq (Batch API) Stacked discounts reach ~25% of list
    High-volume, cost-capped production Fireworks / DeepInfra $0.20 per 1M in + 1M out on identical weights
    Closed frontier models Neither Both serve open weights only
    Sustained 24/7 single-model load Self-host Fixed GPU cost beats per-token above a break-even

    That last row matters more as open weights close the quality gap. We ran the self-hosting break-even in our Qwen3.8-Max open weights breakdown, and the logic holds here.

    Note the hard limit on both vendors: neither serves closed frontier models. If your stack depends on the newest proprietary coder, this comparison does not apply — see our GLM-5.3 vs DeepSeek V4 Pro comparison for the open-weight options that do run here.

    What do the financials say about who wins?

    Cerebras has the better technology and the shakier income statement. It went public on May 14, 2026, raising $5.5 billion. The stock priced at $185, opened at $385 for a 108% pop, and closed at $311 — a $66 billion valuation. Three months later the market repriced it hard.

    The Q2 miss

    On August 12, 2026, Cerebras reported Q2 revenue of $180.1 million against a $193.6 million consensus, per Investing.com. Adjusted EPS came in at a $2.98 loss versus an expected $0.18 loss. Shares fell 14% after hours.

    Core revenue still grew 103% year over year to $209.9 million, and the company guided FY2026 core revenue to $880–890 million. Growth is not the problem. Margin is: guided core operating margin sits at negative 19% to negative 17%.

    Why that should affect your buying decision

    A vendor losing money on every wafer has two exits: raise prices or get acquired. Cerebras’ 2025 revenue was $510 million on 76% growth with $237.8 million of net income, so the balance sheet is not fragile — but the 2026 trajectory is being funded, not earned.

    TrendForce values the company’s three-year OpenAI partnership at over $20 billion. That is concentration risk dressed as a moat.

    The broader point TrendForce makes is the one to internalize: inference is a recurring cost tied directly to revenue, while training is a one-time R&D expense. Its example is brutal — Taalas’ HC1 delivers Llama 3.1 8B at 0.75 cents per million tokens against 3.79 cents on an Nvidia B200.

    Specialized silicon is roughly five times cheaper per token than general-purpose GPUs at that scale. That is why the price you lock in today is unlikely to be the price in twelve months.

    Frequently asked questions about Cerebras vs Groq

    Is Cerebras faster than Groq?

    Yes. Artificial Analysis measures Cerebras at 1,641 tokens per second on GPT-OSS-120B; Groq’s published figure for the same model is around 500. Cerebras also leads on time to first token at 0.46 seconds.

    Is Groq cheaper than Cerebras?

    Yes. On GPT-OSS-120B, Groq charges $0.15 input and $0.60 output per million tokens versus Cerebras at $0.35 and $0.75. With Groq’s stacked batch and cache discounts, the effective gap widens sharply.

    Does Nvidia own Groq now?

    Nvidia agreed in December 2025 to acquire Groq’s assets for about $20 billion, CNBC reported. The Groq API continues to operate and continues to publish its own pricing.

    Can I run Claude or GPT-5 on Cerebras or Groq?

    No. Both providers serve open-weights models only — GPT-OSS, Llama, Qwen, Gemma and similar. Closed frontier models stay on their vendors’ own APIs.

    When does the CS-4 speed actually arrive for API users?

    Cerebras says first CS-4 shipments began in Q3 2026. The 4,400 tokens per second per user figure is a vendor claim on new hardware, not yet an independently measured API result.

    Is paying for faster inference ever worth it?

    Only when latency is visible to a customer or compounds across an agent loop. At roughly 11 cents per minute of latency removed, speed is cheap for interactive products and pure waste for background jobs. We covered the same trade at the frontier-model layer in OpenAI Ultrafast vs Claude Fast Mode.

    The bottom line

    Route by whether something is waiting. If a human or an agent loop blocks on the token stream, Cerebras is worth its 25% output premium — 3.3x measured throughput for that price is one of the better deals in AI infrastructure, and the CS-4 should widen it.

    If nothing is waiting, Cerebras is a rounding-error upgrade you are overpaying for. Send batch and evaluation traffic to Groq’s Batch API at roughly a quarter of list, or to Fireworks and DeepInfra at $0.20 per million in and out.

    The strategic read is less comfortable. Nvidia owns Groq’s technology, Cerebras is losing money at negative 17% to 19% core operating margin, and specialized silicon is already showing five-fold cost advantages per token. Sign nothing longer than twelve months.

    Sources

  • Nvidia Cuts Its OpenAI Data Center Guarantee From $250B to $120B

    Nvidia has cut its financing guarantee for OpenAI’s planned Ohio data center from $250 billion to under $120 billion, according to The Wall Street Journal. The revised backstop covers roughly the first 5 gigawatts of a 10-gigawatt campus. At the same time, Nvidia is in talks to invest up to $3 billion in SB Energy, the SoftBank unit building it.

    The Nvidia OpenAI data center guarantee is now roughly half what it was three weeks ago. Nothing about the physical project changed. What changed was how much risk Nvidia’s own shareholders were willing to let the company carry.

    That distinction matters more than the headline number.

    What exactly did Nvidia change?

    Nvidia reduced the credit guarantee it would provide behind OpenAI’s lease of the Ohio campus. The Wall Street Journal reported the figure fell from up to $250 billion to less than $120 billion. The smaller backstop now covers only about the first 5 gigawatts of the 10-gigawatt site.

    A backstop is not cash. It is a promise: if OpenAI cannot pay the lease, Nvidia does.

    That promise is what makes the project financeable. Lenders will not underwrite a $500 billion buildout against an unprofitable tenant. They will underwrite it against Nvidia’s balance sheet.

    The numbers, before and after

    Item Reported July 27, 2026 Reported August 14–15, 2026
    Lease guarantee from Nvidia Up to $250 billion Under $120 billion
    Capacity covered Full 10 GW campus First ~5 GW
    Separate chip financing discussed ~$350 billion Not restated
    Nvidia equity stake in SB Energy Not discussed Up to $3 billion, in talks
    Status of OpenAI lease In negotiation Still not binding

    Reuters reported that OpenAI has still not signed a binding lease for the full project. That is worth holding onto. Every figure above describes a deal that does not yet legally exist.

    Why did Nvidia scale the guarantee back?

    Investors pushed back. According to the Journal’s reporting, the change followed concerns about Nvidia’s risk exposure tied to very large financing commitments on projects that are not yet operating. The company trimmed the obligation rather than defend it.

    This is the part worth pausing on.

    Nvidia’s fiscal 2026 revenue was $215.9 billion with net income of $117 billion, and it held $62.6 billion in cash and equivalents as of January 25. A $250 billion contingent obligation is larger than the company’s entire annual revenue. Halving it does not make it small.

    The circular-financing problem nobody has solved

    Nvidia sells chips. Nvidia also funds the companies that buy the chips. As The Next Web noted, Nvidia spent more than $40 billion on equity positions in the first four months of 2026, and almost all of it went to firms that purchase its hardware.

    Nvidia’s Q2 2026 13F filing, submitted August 14, showed 122.8 million SpaceX Class A shares worth roughly $21 billion and 214.8 million Intel shares worth about $30 billion, the latter built from an initial $5 billion investment.

    Revenue that depends on capital you supplied is not the same quality of revenue as a customer paying from their own cash flow. That is the honest read, and it applies whether the guarantee is $250 billion or $120 billion.

    What is the Ohio data center campus?

    SB Energy, a SoftBank Group company, is developing a 10-gigawatt campus at the Portsmouth site in Pike County, Ohio, on federal land owned by the US Department of Energy. Data Center Dynamics reports a first phase of roughly 800 megawatts targeted to begin operating in 2028.

    Full build-out is estimated at $500 billion. Ground was broken in March 2026.

    If completed, it would be the largest data center project ever announced.

    Power, not silicon, is the binding constraint

    The energy math is the story underneath the story. The project requires roughly 9.2 gigawatts of new natural gas generation, plus about $4.2 billion of transmission work with AEP Ohio, according to reporting on the plan.

    Chips arrive in months. Gas turbines and transmission lines take years.

    • 10 GW — total planned campus capacity
    • 800 MW — first phase, targeted for 2028
    • 9.2 GW — new gas generation required
    • $4.2 billion — transmission work with AEP Ohio
    • $500 billion — estimated cost at full build-out

    This is why the money is moving toward power developers rather than pure compute. It is the same shift that has been reshaping the largest corporate capex commitments in AI.

    Why is Nvidia buying a stake in SB Energy?

    The Information reported that Nvidia is negotiating an investment of up to $3 billion in SB Energy, structured roughly 50/50: about $1.5 billion at signing, the rest tied to SB Energy’s planned IPO. Goldman Sachs is advising SB Energy; Morgan Stanley is advising Nvidia.

    SB Energy could go public as soon as September 2026, seeking to raise at least $5 billion.

    Read the two moves together and a pattern appears. Nvidia is swapping an open-ended contingent liability for a defined equity position — less downside exposure, more upside participation.

    The timing is not an accident

    Trimming a guarantee weeks before your partner’s IPO is a signal to public-market buyers about how much of the project’s credit risk sits with a third party. A cleaner structure is easier to price.

    Whether it is easier to sell is a different question. SoftBank carries more than $130 billion in debt.

    Who profits from this?

    Nvidia announced partnerships with six major financial institutions this week to build compute financing platforms, part of an effort to mobilize more than $500 billion in third-party capital for AI infrastructure. The direction of travel is clear: move the risk off Nvidia’s books and onto someone else’s.

    Banks earn fees. SoftBank monetizes an asset. Utilities and gas turbine makers get multi-year order books.

    OpenAI, valued at $852 billion after its record $122 billion raise in March 2026, gets compute it could not finance alone — while remaining unprofitable, with projected compute spending of roughly $750 billion through 2030.

    Why this matters

    The AI trade has quietly become a credit trade. The bottleneck is no longer model quality or chip supply; it is who will underwrite twelve-figure obligations against tenants that do not yet generate profit.

    When the largest supplier in the industry halves its own guarantee under shareholder pressure, that is a data point about the market’s appetite for that risk. It is not a collapse. It is a repricing.

    Watch three things: whether the binding lease is signed, whether SB Energy’s IPO clears at target size, and whether other vendors follow Nvidia in shifting from guarantees to equity. Similar structural pressure is visible across the global chip supply chain and in how private AI companies such as Databricks and Anthropic are raising capital.

    This article is reporting and analysis, not financial advice.

    Frequently asked questions

    How much did Nvidia cut the OpenAI data center guarantee?

    From up to $250 billion down to less than $120 billion, per The Wall Street Journal. The revised amount covers roughly the first 5 gigawatts of the planned 10-gigawatt Ohio campus.

    Is the OpenAI Ohio lease signed?

    No. Reuters reported that OpenAI was still negotiating a binding lease for the full project as of mid-August 2026. Reports suggested a signing could come as soon as that weekend.

    What is SB Energy?

    SB Energy is a SoftBank Group company founded in 2019 that develops power generation and data center campuses. OpenAI and SoftBank each invested $500 million in it in January 2026.

    When is the SB Energy IPO?

    Reports indicate SB Energy could list as soon as September 2026, targeting a raise of at least $5 billion. No prospectus terms have been confirmed publicly.

    Why does a chipmaker guarantee a lease at all?

    Because lenders will not finance a $500 billion project against an unprofitable tenant. Nvidia’s credit makes the debt cheaper, which accelerates construction and, ultimately, chip orders.

    What is circular financing in AI?

    It describes vendors funding their own customers. Nvidia spent over $40 billion on equity in early 2026, largely in companies that buy its hardware, which makes some of its revenue partly self-financed.

    How big is the Ohio project compared with others?

    At 10 gigawatts and an estimated $500 billion, it would be the largest data center project announced to date if completed. The first 800-megawatt phase is targeted for 2028.

    The bottom line

    Nvidia did not walk away. It renegotiated its exposure downward by more than $130 billion and replaced part of it with an equity stake it can sell.

    That is a rational trade for Nvidia. It is a harder one for everyone downstream, because the capital that Nvidia stopped guaranteeing has to come from somewhere — banks, bond markets, or public IPO buyers who will price the risk more honestly than a vendor guarantee ever did.

    The next two data points are the binding lease and the SB Energy listing. If both land on schedule, the buildout continues on cheaper terms. If either slips, the market will learn what a 10-gigawatt campus is worth without a chipmaker’s signature behind it.

    Sources