Author: wealthenginex

  • Anthropic Music Copyright Lawsuit Puts $150,000 a Song on the Table

    Sony Music Publishing and Warner Chappell have filed an Anthropic music copyright lawsuit in the US District Court for the Northern District of California, naming CEO Dario Amodei and co-founder Benjamin Mann as defendants. The complaint covers “tens of thousands” of compositions and seeks up to $150,000 per willfully infringed work plus $25,000 per stripped copyright notice — a multi-billion-dollar claim landing weeks before Anthropic’s expected IPO.

    What does the Anthropic music copyright lawsuit actually claim?

    The publishers accuse Anthropic of building Claude on pirated material. The complaint alleges “a brazen campaign of illegally torrenting, scraping, and downloading” copyrighted works, according to Axios. It calls the conduct one of the largest ongoing thefts of intellectual property in history.

    Music Business Worldwide dates the filing to August 28; TechCrunch and Axios published their reports on August 29.

    Three sourcing channels are alleged. Torrented book libraries. Scraped lyric sites. And bulk web datasets.

    Where the training data allegedly came from

    • Library Genesis — Mann allegedly downloaded at least 5 million pirated books in June 2021, per Music Business Worldwide.
    • Pirate Library Mirror — employees allegedly torrented roughly 2 million further works in July 2022, the same report says.
    • Licensed lyric services — the complaint alleges scraping from MusixMatch and LyricFind.
    • Bulk datasets — Common Crawl, The Pile and Books3.
    • Second-hand book scanning — physical acquisition and digitization operations.

    Books matter here because they carry lyrics and sheet music. That is how a literary-piracy allegation becomes a music-publishing claim.

    Why the founders are named personally

    Naming Amodei and Mann is a pressure tactic as much as a legal one. Individual defendants complicate insurance, discovery and settlement, and they make depositions personal. Anthropic has fought this before: its CEO moved to drop a direct-infringement claim against him in the earlier publishers’ case, MBW reported.

    How much money is actually at risk?

    The honest answer is that nobody knows, because the “multi-billion” figure is a statutory ceiling multiplied by an unspecified work count. Statutory damages run to $150,000 per willfully infringed work, plus $25,000 for each removal of copyright management information, according to Engadget.

    Item Figure Source
    Statutory ceiling, willful infringement Up to $150,000 per work Engadget / MBW
    Ceiling for removing copyright management info Up to $25,000 per violation MBW
    Works alleged in this case “Tens of thousands” of compositions Axios
    Earlier publishers’ suit (UMG, Concord, ABKCO) $3B sought over 20,000+ songs TechCrunch, Jan 2026
    Authors’ settlement $1.5B Axios
    Anthropic annualized revenue $65B+ (end of July 2026) TechCrunch / Bloomberg

    Run the arithmetic on the January case and the anchoring becomes obvious. Twenty thousand songs at $150,000 each is exactly $3 billion. Plaintiffs are pricing at the statutory maximum, which is a negotiating position, not a forecast.

    Courts rarely award the ceiling. The authors’ matter resolved at $1.5 billion — a large number, but one produced by settlement rather than a maxed-out jury verdict.

    There is a second reason to discount the headline. “Tens of thousands” is not a work count a court can rule on. Until the publishers file an exhibit listing every composition, registration by registration, the exposure is a range, not a number.

    Registration status matters too. Statutory damages generally require timely US registration, and catalogs of this size rarely qualify uniformly. Expect the defense to attack the eligible-work count long before it argues fair use.

    Why does the timing matter so much?

    Because Anthropic is trying to go public. The company has filed confidential IPO paperwork and could reach the market as soon as this fall, seeking a public valuation of $2 trillion or more, TechCrunch reported on August 17. Unquantified litigation is exactly what underwriters hate.

    The financial backdrop is strong. Annualized revenue passed $65 billion at the end of July, up from $47 billion in May and $9 billion at the end of 2025, per the same report. The May round valued the company at $965 billion.

    Against $65 billion of annualized revenue, even a $3 billion judgment is survivable. Against an S-1 risk-factors section, it is a paragraph every institutional buyer will read twice.

    The pattern in the docket

    1. June 2021 / July 2022 — the alleged torrenting activity now cited across multiple complaints.
    2. January 29, 2026 — UMG, Concord and ABKCO sue for $3 billion over 20,000+ songs.
    3. 2026 — the $1.5 billion authors’ settlement clears court approval, per MBW.
    4. August 28, 2026 — Sony Music Publishing and Warner Chappell file.

    Each case reuses evidence surfaced by the last one. That compounding discovery record is the real liability, not any single filing.

    Who profits from this?

    Music publishers, first. Every settlement resets the market price of training data and converts back catalogs into recurring licensing revenue. Litigation is functioning as price discovery for an asset class that had no clearing price two years ago.

    The second beneficiary is the incumbent AI lab with cash. Licensed data is expensive, and expense favors scale. A challenger training on scraped corpora now inherits a liability that a well-capitalized lab can simply buy its way out of.

    The losers are mid-size labs and open-weight projects that cannot write nine-figure checks to rights holders.

    That consolidation effect is underrated. Copyright enforcement is often framed as a check on big AI companies, but the practical result is a moat: only firms with $65 billion revenue run rates can absorb the licensing bill and the legal reserve at the same time.

    Why this matters for the wider AI market

    Training-data liability has moved from a legal footnote to a balance-sheet line item. The prior ruling in the authors’ litigation drew the line clearly: using copyrighted works for training could be lawful, but acquiring them through piracy was not, TechCrunch noted.

    That distinction is the whole ballgame. It shifts the fight from “is AI training fair use” — an argument labs have often won — to “how did you get the files,” which leaves a forensic trail.

    For investors reading the AI infrastructure trade, this is a cost input alongside compute. Anthropic has committed enormous sums to capacity, including the $45 billion Nscale datacenter deal and the Broadcom-linked chip financing package. Data licensing is becoming a third structural cost next to silicon and power.

    There is a precedent risk beyond Anthropic. The same acquisition-versus-use distinction applies to every lab that touched Books3 or Library Genesis, and the same music publishers hold catalogs they can assert repeatedly. One favorable ruling here becomes a template.

    It also complicates brand positioning. The complaint leans hard on the gap between Anthropic’s safety-first marketing and its alleged sourcing — a reputational angle the company has faced before, as during the Claude watermark backlash.

    This post is reporting and analysis, not financial advice.

    Frequently asked questions

    Who filed the Anthropic music copyright lawsuit?

    Sony Music Publishing and Warner Chappell Music, in the US District Court for the Northern District of California. Anthropic, Dario Amodei and Benjamin Mann are named as defendants.

    How many works are involved?

    The complaint refers to “tens of thousands” of copyrighted compositions, per Axios — broader than the 20,000-plus songs at issue in the January 2026 publishers’ case.

    What damages are the publishers seeking?

    Up to $150,000 per willfully infringed work and up to $25,000 per removal of copyright management information, which the plaintiffs say totals billions.

    Has Anthropic responded?

    Not at the time the initial reports published. TechCrunch, Axios and Engadget all noted the company had not provided comment.

    Is this the same as the earlier music lawsuit?

    No. Universal Music Group, Concord and ABKCO filed a separate $3 billion action in January 2026. BMG has its own narrower case. This is a new, larger filing by two different publishers.

    Does this affect Anthropic’s IPO?

    It adds a material, unquantified risk factor. Anthropic has filed confidentially and is reportedly targeting a public valuation above $2 trillion, so the disclosure will be scrutinized closely.

    What else is the lawsuit asking for?

    Beyond damages, the publishers seek destruction of infringing copies, a full accounting of Claude’s training data and a jury trial, according to Music Business Worldwide.

    The bottom line

    The Anthropic music copyright lawsuit is unlikely to be decided on its headline number. Statutory maximums are an opening bid, and the prior authors’ matter shows how far a settlement can land from the ceiling.

    What matters is the discovery record. Each successive complaint borrows evidence from the last, and the accounting of Claude’s training data that the publishers demand would be far more damaging than any single payout.

    Watch three things: whether Anthropic settles before pricing its IPO, whether the founders stay named as defendants, and whether the S-1 quantifies the exposure. The first would be the clearest signal that the company wants this off the table before it faces public-market investors.

    Sources

  • Wan 3.0 vs Veo 3.1: Which 30-Second AI Video Costs Less

    Wan 3.0 vs Veo 3.1 comes down to one question: do you need 30 seconds in a single pass, or four eight-second cuts? Wan 3.0 delivers a continuous 30-second 1080p clip with audio for $6.00. Veo 3.1 Fast delivers the same length for $3.60 — but in four separate generations you have to stitch. Wan wins continuity; Veo Fast wins the invoice.

    What is Wan 3.0, and what changed on August 24?

    Alibaba formally launched Wan 3.0 on August 24, 2026, after a public beta that opened on August 6. The headline change is duration: 30 seconds in one generation, double the 15-second ceiling of Wan 2.7, with audio produced in the same pass.

    That single number reorders the market. Until now, every long AI video was a stitching job.

    The model takes text, images, video, audio, documents and public web pages as input. Alibaba’s spec allows up to 10 images, 5 video clips and 5 audio tracks in one request, plus documents up to 50 pages or 100 MB, according to fal’s model page.

    The specs that matter for cost

    Output runs at 30 fps in 480p, 720p or 1080p, with 1080p as the default. Duration is either set manually from 2 to 30 seconds or chosen by the model. Audio — ambient sound, effects and cross-lingual lip sync — is on by default rather than bolted on afterward.

    Scenario’s rundown also documents a six-shot “AI Director” mode that generates multi-shot sequences with consistent characters from one prompt. Independent testing by Atlas Cloud found unwanted cuts inside shots that were meant to stay continuous, so the single-pass claim is not yet flawless.

    How much does Wan 3.0 cost per video?

    Wan 3.0 bills per second of output, scaled by resolution: $0.05/s at 480p, $0.10/s at 720p and $0.20/s at 1080p. A full 30-second generation therefore lands between $1.50 and $6.00. Audio is included at every tier — there is no separate sound charge.

    Here is how that compares against the rest of the field on official per-second rates, normalized by invideo’s August 2026 pricing survey.

    ModelOfficial rateMax single passNative audio30s at best quality
    Wan 3.0$0.20/s (1080p)30sYes$6.00 — 1 generation
    Veo 3.1$0.40/s (720p–1080p)8sYes$12.00 — 4 generations
    Veo 3.1 Fast$0.12/s (1080p)8sYes$3.60 — 4 generations
    Veo 3.1 Lite$0.05/s (720p)8sYes$1.50 — 4 generations
    Seedance 2.5~$0.14/s (1080p, token-billed)30sYes~$4.20 — 1 generation
    MiniMax-H3$0.13/s (2K)15sYes (stereo)$3.90 — 2 generations
    Kling 3.0 Turbo~$0.11/s (720p)15sYes~$3.30 — 2 generations
    sora-2-pro$0.70/s (1080p)20sYes$21.00 — API sunsets Sept 24
    Per-second rates as published August 2026; 30-second totals are arithmetic on those rates.

    Note the free tier: fal gives signed-in users five Wan 3 text-to-video generations every 24 hours at 720p with audio. That is roughly $15 of daily inference at list price, which makes evaluation cheap.

    Which is cheaper for a 30-second clip, Wan 3.0 or Veo 3.1?

    Veo 3.1 Fast is cheaper on paper — $3.60 versus $6.00 for 30 seconds at 1080p, a 40% saving. But Veo 3.1 caps a single generation at 4 to 8 seconds. Hitting 30 seconds means four generations, four prompt-consistency gambles, and an edit.

    Against standard Veo 3.1 at $0.40/s, Wan 3.0 is 50% cheaper for the same length. Against sora-2-pro at $0.70/s, it is 71% cheaper — and Sora’s API sunsets on September 24, 2026, so that comparison has a shelf life. We covered the migration options in our Sora 2 alternatives breakdown.

    The stitching tax nobody puts in the pricing table

    Extension chains are not free quality. Veo 3.1 can reach 148 seconds through up to 20 seven-second extensions, but that path is 720p-only. Sora 2’s API tops out at 120 seconds across six extensions.

    Every splice is a place where lighting, face and wardrobe drift. The invideo length survey, dated August 7, 2026, recommends assembling separate short shots rather than trusting long extension chains, precisely because degradation accumulates.

    Wan 3.0’s counter-argument is that one pass has no splices. Its cost argument is weaker than its workflow argument.

    Where Wan 3.0’s pricing actually hurts

    Retries. A rejected 30-second 1080p Wan generation burns $6.00. A rejected 8-second Veo 3.1 Fast shot burns $0.96. If your acceptance rate on the first try is 50%, Wan’s effective cost per delivered 30 seconds is $12.00 — above standard Veo 3.1 and well above the Fast tier.

    That is the number to model before committing a production pipeline. Long single-pass generation concentrates risk in one expensive request.

    Is Wan 3.0 the best AI video model right now?

    No — not on the evidence available. Wan 3.0 has no independent arena score yet. In the Artificial Analysis blind-vote text-to-video-with-audio arena for August 2026, the top three are Gemini Omni Flash at 1245 Elo, MiniMax-H3 at 1242 and Seedance 2.0 at 1225.

    Wan 2.7 — the predecessor, not the new model — sits around 1163. Kling 3.0 scores 1113. Veo 3.1, despite being the default enterprise choice, ranks eleventh at 1098.

    What the arena actually measures

    Blind human votes on short prompts. It rewards visual polish and audio sync; it does not measure 30-second continuity, document ingestion or multi-shot direction — the three things Wan 3.0 was built for.

    So treat 1163 as a floor for Wan 3.0’s likely placement, not a verdict. And note that the arena leader, Gemini Omni Flash, bills in tokens rather than seconds: 5,792 tokens per second of 720p video, per Google’s pricing docs. Token billing makes budget forecasting harder, which is a real cost.

    Which AI video model should you use for each job?

    Pick by shot length and failure tolerance, not by benchmark rank. If your deliverable is one continuous scene longer than 15 seconds, the field narrows to two models. If it is a stack of social cuts, the cheapest per-second tier wins outright.

    JobPickWhyBudget for 30s
    One continuous 30s scene, 1080pSeedance 2.5Only single-pass 30s option that undercuts Wan~$4.20
    30s scene from a PDF or deckWan 3.0Document-to-video, 50 pages / 100 MB$6.00
    Batch of 8s social cutsVeo 3.1 Fast$0.12/s at 1080p, cheap retries$3.60
    Highest blind-vote qualityMiniMax-H31242 Elo, native 2K, stereo audio$3.90
    Cheapest usable draftWan 3.0 at 480p$0.05/s with audio, full 30s$1.50
    Multi-shot narrativeKling 3.0Six labeled shots per generation~$3.30
    New builds on SoraMigrate nowAPI sunsets September 24, 2026n/a
    Prices derived from published per-second rates, August 2026.

    Is Wan 3.0 open weights?

    No. Despite third-party pages advertising Apache 2.0 licensing, Wan 3.0 shipped as a hosted API with no Hugging Face weights, no GitHub repository and no published license as of late August 2026, per Atlas Cloud’s audit.

    Wan 2.7 is closed too. The last downloadable Wan flagship is Wan 2.2, released July 28, 2025 under Apache 2.0.

    That matters financially. Self-hosting is the only route to sub-cent-per-second video, and it is closed here. Alibaba’s own docs ask for at least 80 GB of VRAM for single-GPU A14B inference; the smaller TI2V-5B variant runs on 24 GB consumer cards. Compare that with the text-model side, where Tencent’s Hy4 release put 770B open weights in play.

    What does this mean for your video budget?

    Three shifts follow directly from the August 24 launch. None of them require you to switch providers this week, but all of them change the math on annual spend.

    • Per-second rates are converging. Wan 3.0 at 720p costs $0.10/s — identical to Veo 3.1 Fast at 720p. Differentiation has moved from price to clip length.
    • Duration is the new premium feature. Only Seedance 2.5 and Wan 3.0 hit 30 seconds natively. Expect that capability to carry a markup until Veo and Kling match it.
    • Retry rate now dominates unit cost. At 1080p, one rejected Wan generation costs more than six Veo Fast shots. Track your first-pass acceptance rate before scaling.
    • Closed weights cap the downside. With no self-host path for Wan 3.0 or 2.7, list price is your price. Budget accordingly.

    Frequently asked questions

    How much does a 30-second Wan 3.0 video cost?

    $1.50 at 480p, $3.00 at 720p and $6.00 at 1080p, with audio included at every tier. Billing is per second of output.

    Can Veo 3.1 generate 30 seconds in one shot?

    No. Veo 3.1 caps a single generation at 4 to 8 seconds. Longer runs require extension chaining, which reaches 148 seconds but only at 720p.

    Is Wan 3.0 better than Sora 2?

    On cost, clearly: $6.00 versus $21.00 for 30 seconds at 1080p on sora-2-pro. On availability, decisively — Sora’s API sunsets September 24, 2026.

    Which AI video model ranks highest right now?

    Gemini Omni Flash leads the Artificial Analysis text-to-video-with-audio arena at 1245 Elo, ahead of MiniMax-H3 at 1242 and Seedance 2.0 at 1225. Wan 3.0 has no independent score yet.

    Does Wan 3.0 generate audio automatically?

    Yes. Audio is on by default and generated in the same pass as the video, including ambient sound, effects and cross-lingual lip sync.

    Can I run Wan 3.0 on my own GPU?

    Not currently. No weights have been published. The last downloadable Wan flagship is Wan 2.2 from July 2025, and single-GPU A14B inference is documented as needing 80 GB of VRAM.

    What is the cheapest way to make a 30-second AI video with audio?

    Wan 3.0 at 480p, at $1.50 for the full clip in one generation. For 1080p output, Veo 3.1 Fast at $3.60 is cheapest if you can accept four stitched eight-second shots.

    The bottom line

    Buy Wan 3.0 for continuity, not for price. If your deliverable is a single unbroken 30-second scene at 1080p with synced audio — a product film, an explainer, a narrative ad — it is the strongest option at $6.00, and the document-to-video input has no equivalent elsewhere.

    If you are producing volume short-form, do not switch. Veo 3.1 Fast at $0.12/s and Kling 3.0 Turbo at roughly $0.11/s remain cheaper per delivered second, and their smaller generations make retries painless.

    If you want maximum quality per dollar and can live with 15 seconds, MiniMax-H3 at $0.13/s for native 2K is the value pick — it ranks second on the arena while Veo 3.1 ranks eleventh.

    The one clear loser is Sora. With the API sunsetting September 24 and sora-2-pro priced at $0.70/s, there is no financial case for building anything new on it. For the equivalent decision on still images, see our AI image generation API rankings.

    Sources

  • Tencent Hy4: 770B Open Weights, 82% Cheaper Output Than Kimi K3

    Tencent Hy4 preview is a 770-billion-parameter open-weight model released under Apache 2.0 on August 28, 2026, with 49B active parameters and a 1,048,576-token context window. Tencent Cloud prices it at $0.834 per million input tokens and $2.501 per million output — roughly one-eighth the output cost of GPT-5.6 Sol and 82% below Kimi K3. Weights are on Hugging Face.

    Tencent open-sourced its largest model to date on Friday, and the interesting number is not the parameter count. It is the price tag attached to weights anyone can download, modify and resell.

    That combination — frontier-adjacent scores, permissive licensing, and output tokens at $2.501 per million — is the part that moves money.

    What is Tencent Hy4 preview?

    Tencent Hy4 preview is a mixture-of-experts language model with 770B total parameters, of which 49B activate per token. Tencent released it on August 28, 2026, published the weights on Hugging Face under the Apache 2.0 license, and shipped it simultaneously into its own consumer and developer products.

    According to Tencent’s announcement, the model is live in WorkBuddy, CodeBuddy, Yuanbao and ima, with API access through Tencent Cloud TokenHub and OpenRouter.

    WorkBuddy and CodeBuddy are free for two weeks from launch. That is a customer-acquisition subsidy, not a pricing model.

    The architecture, briefly

    The Hugging Face model card lists 78 layers — the first a dense FFN, the remaining 77 MoE — with 256 routed experts plus one shared expert per MoE layer, and top-8 routing per token.

    Vocabulary size is 120,832. Attention is what Tencent calls “Gated Sparse Attention with IndexCache,” reusing sparse indices across layers to keep the million-token window affordable to serve.

    There is also a built-in multi-token-prediction layer of 10B parameters (0.7B activated) for speculative decoding. Tencent says system-level optimization lifted end-to-end throughput 31.8% against its own baseline.

    How much does Tencent Hy4 cost?

    Tencent Cloud lists $0.834 per million input tokens, $2.501 per million output tokens, and $0.042 per million on cache hits. In renminbi terms Tencent quotes ¥6 input and ¥18 output per million tokens.

    Those are the headline numbers, and they are aggressive against every comparable model.

    Per KuCoin’s summary of the launch materials, that input price is 25% below GLM-5.3 and 70% below Kimi K3; on output it is 36% below GLM-5.3 and 82% below Kimi K3.

    Model Input / 1M Output / 1M Context Weights
    Tencent Hy4 preview $0.834 $2.501 1,048,576 Apache 2.0
    GPT-5.6 Sol (base tier) $4.00 $20.00 n/d Closed
    GPT-5.6 Sol (long context) $8.00 $30.00 n/d Closed

    On output — the expensive half of any agentic workload, where the model writes code, calls tools and re-reads its own work — Hy4 runs at roughly one-eighth of GPT-5.6 Sol’s base rate.

    For a coding agent burning 50 million output tokens a month, that is the difference between about $1,000 and about $125. The comparison only holds if the cheaper model finishes the job in a similar number of tokens, which is exactly the assumption worth testing.

    How does Hy4 score on benchmarks?

    Hy4 posts strong software-engineering numbers and slightly trails the closed frontier on reasoning. Tencent reports 92.3 on GPQA Diamond, 65.7 on SWE-Bench Pro, 82.9 on SWE-Bench Multilingual, 62.9 on SkillsBench v1.1, and 64.3 on Deep-SWE.

    The Deep-SWE result is the standout: 64.3 against 28.0 for the previous generation. That is not incremental.

    Against closed models the gap is real but narrow. Hy4 scores 92.3 on GPQA Diamond versus GPT-5.6 Sol’s 94.6, and 85.4 on Terminal-Bench versus Sol’s 88.8.

    • GPQA Diamond: Hy4 92.3 — GPT-5.6 Sol 94.6
    • Terminal-Bench: Hy4 85.4 — GPT-5.6 Sol 88.8
    • SWE-Bench Multilingual: Hy4 82.9
    • SWE-Bench Pro: Hy4 65.7
    • Deep-SWE: Hy4 64.3, up from 28.0

    Every one of those figures is vendor-reported. None has been independently reproduced at the time of writing, and the model has been public for about a day.

    Is Hy4 better than GLM-5.3 and Kimi K3?

    Marginally, on Tencent’s own evidence. Tencent ran a blind evaluation with 163 internal experts across 203 real engineering tasks. Hy4 preview averaged 2.99 out of 4.00, against 2.94 for Kimi K3 and 2.92 for GLM-5.3.

    A 0.05-point spread on a four-point scale is not a capability gap. It is a tie with a favorable rounding.

    The win-rate breakdown is more honest about how close this is. Against GLM-5.3, Hy4 won 46.8% of comparisons, drew 12.8% and lost 40.4%. Against Kimi K3: 51.2% wins, 7.9% draws, 40.9% losses.

    So Hy4 loses roughly two of every five head-to-head comparisons against models that were already open. And per KuCoin’s read of the same materials, Hy4 “did not lead comprehensively in public benchmarks” and lags GLM-5.3 in code and cybersecurity tests.

    The differentiator here is price and license, not raw capability. That is worth saying plainly, because the launch framing does not. If you are choosing between cheap open coders, our GLM-5.3 vs DeepSeek V4 Pro comparison and the GLM-5.3-Flash vs Qwen3.8-Flash-Next breakdown cover the alternatives.

    Should you self-host Hy4 or use the API?

    Self-hosting is viable at a scale that would have required a cluster a year ago. Tencent documents deployment on a single eight-GPU node using the FP8 quantized variant, with prebuilt containers for vLLM and SGLang at tensor-parallel size 8.

    The full BF16 checkpoint and an FP8 build are both published, along with the AngelSlim toolkit for further compression.

    The API case is stronger for anyone under roughly 100 million tokens a month. One eight-GPU node of current-generation accelerators plus the engineer who babysits it will not come in under $2.501 per million output tokens at low volume.

    The self-host case is stronger for three groups: teams with data-residency constraints, teams already running GPU capacity at low utilization, and teams that want to distill Hy4 into something smaller. Apache 2.0 permits all three without a negotiation. Our earlier analysis of Qwen3.8-Max open weights versus API walks through that math in detail.

    Who wins and loses financially?

    The clearest loser is anyone selling mid-tier closed inference. If a 770B open model at $2.501 output holds up in production, the price umbrella over $20-per-million output tiers gets thinner.

    Kimi K3 is the most directly exposed. An 82% output-price gap against a model that wins only 51.2% of blind comparisons is a hard position to hold.

    Tencent wins on distribution, not on API margin. At ¥6 per million input, the API is a loss leader that routes developers toward Tencent Cloud, WorkBuddy and CodeBuddy — the same playbook that made cheap Chinese inference a strategic instrument rather than a business line. We covered the last turn of that cycle when DeepSeek raised prices by up to 1,100%.

    The GPU vendors win either way. Open weights that need eight accelerators per node create hardware demand that a closed API never surfaces on anyone else’s balance sheet.

    What are the catches?

    Three, and Tencent names one of them itself.

    The company acknowledges in its release notes that the model spends “longer than necessary reasoning” and over-verifies its own work. On a consumption-priced endpoint, verbosity is a bill. A model that is 82% cheaper per token but writes twice as many tokens is 64% cheaper, not 82%.

    Second, serving is thin. OpenRouter shows a single provider — Tencent Cloud — with 3.16-second P50 latency, 38 tokens per second, and 98.65% availability over three days. There is no failover route.

    Third, the context window is a spec, not a guarantee. The model accepts 1,048,576 input tokens but caps completions at 64,000, and nothing in the release claims uniform recall across the full window. For a like-for-like look at long-context pricing, see our cheapest 1M-context model comparison.

    And it is called “preview” for a reason.

    Frequently asked questions

    Is Tencent Hy4 preview free to use?

    The weights are free under Apache 2.0. The hosted API is not — it costs $0.834 per million input tokens and $2.501 per million output. WorkBuddy and CodeBuddy are free for two weeks from the August 28 launch.

    Can I use Hy4 commercially?

    Yes. Apache 2.0 permits commercial deployment, modification, distillation and redistribution without a separate license negotiation or revenue threshold.

    What hardware do I need to run Hy4?

    Tencent documents a single eight-GPU node using the FP8 quantized build, served through vLLM or SGLang at tensor-parallel size 8. Minimum memory figures are not published.

    How big is the context window?

    1,048,576 tokens of input, with completions capped at 64,000 tokens.

    Is Hy4 better than GPT-5.6 Sol?

    Not on published benchmarks. Hy4 scores 92.3 on GPQA Diamond against Sol’s 94.6, and 85.4 on Terminal-Bench against 88.8. It is cheaper by roughly 8x on output.

    Where can I download the weights?

    Hugging Face at tencent/Hy4-preview, with code and deployment instructions on GitHub.

    Have the benchmarks been independently verified?

    No. All published scores are vendor-reported as of August 29, 2026.

    The bottom line

    Hy4 preview is not a capability breakthrough. It wins its own blind evaluation by 0.05 points and loses 40% of head-to-head comparisons against models that were already open-weight.

    It is a pricing event. Tencent shipped near-parity performance under Apache 2.0 at 82% below Kimi K3’s output rate and roughly one-eighth of GPT-5.6 Sol’s, and put the weights on Hugging Face the same day.

    Tencent’s own README calls it “another step change in capability — the largest generation-over-generation gain we’ve measured.” The blind-evaluation table does not support that framing. The invoice does.

    For anyone running high-volume agentic workloads, Hy4 is worth a benchmark run this week — with token-consumption logging turned on, because the verbosity Tencent admits to is where the savings go to die. For anyone selling inference above $20 per million output tokens, the floor moved again.

    Sources

  • a16z Machine Age Fund: $1.1 Billion Bet on AI Hardware

    The a16z Machine Age Fund closed at $1.1 billion on August 28, 2026, and it buys physical things: chips, memory, networking, power gear, cooling, robots and data center real estate. Andreessen Horowitz says hardware now accounts for more than 20% of its deal flow. The timing is not subtle — Nvidia had just posted $96.2 billion in quarterly revenue two days earlier.

    What is the a16z Machine Age Fund?

    The a16z Machine Age Fund is a $1.1 billion vehicle dedicated to the physical layer of artificial intelligence. It invests in chips, memory, networking, storage, data centers, power, cooling and robotics — not software. Andreessen Horowitz announced it on August 28, 2026.

    That is a real break in character. The firm built its name on Marc Andreessen’s 2011 argument that software was eating the world.

    The new fund concedes that software cannot run without something to run on, and that the something is now the bottleneck.

    Who is running the fund

    According to a16z’s own announcement, the fund is backed by general partners Ben Horowitz, Martin Casado, Raghu Raghuram, David Ulevitch and David George.

    SiliconANGLE reports that partner Guido Appenzeller, formerly chief technology officer of Intel’s data center business, is also on the team. Casado and Raghuram both came from VMware.

    That roster is telling. This is an infrastructure operator bench, not a consumer-app bench.

    How much did a16z raise, and where does the money go?

    The fund is $1.1 billion and spans early and growth stage. Its remit runs the full stack — from silicon to the buildings that house it. PitchBook notes the fund targets chips, memory, networking, storage, data centers and robotics in a single mandate.

    Layer What the fund buys Named a16z holdings
    Silicon Processors, memory, custom accelerators Unconventional AI
    Networking Switching and interconnect for AI clusters Nexthop
    Power Solid-state transformers, electrical infrastructure Heron Power (backed 2025)
    Facilities Data center construction and real estate Volta
    Materials Cooling and advanced materials Atoms
    Robotics Autonomous machines, edge AI hardware Mind Robotics, Skydio, Anduril

    a16z has been writing these checks for a while without a dedicated fund. PitchBook records a $500 million Series B for Nexthop AI in March 2026 and a $500 million Series A for Mind Robotics the same month, co-led with Accel.

    Longer-dated positions include Skydio from 2016, Anduril from 2019 and Waymo from 2020.

    Why is a16z betting on AI hardware now?

    Because the supply chain cannot expand fast enough. a16z’s central claim is a growth-rate mismatch: hardware suppliers are structured to grow 20% to 30% a year, while AI infrastructure demand is growing in triple digits.

    The firm put it bluntly in its announcement: “The hardware industry supply side is used to growing 20% to 30% per year at most; not the triple-digit growth that’s needed to catch up with demand.”

    Every rung of that ladder is constrained at once — chips, memory, power, and the physical space to put them in.

    The power math is the real story

    The numbers a16z cites for rack density explain why this became a hardware problem rather than a software one.

    • Compute density: up 28x from H100 configurations to Rubin racks, per a16z.
    • Power per rack: from 5–10 kW historically to 100–250 kW today, with a16z projecting 1 megawatt per rack within three years.
    • Campus scale: from tens of megawatts to hundreds, with some sites now planned at gigawatt scale.
    • Deal flow shift: hardware has gone from a marginal share of a16z’s pipeline to more than 20%.

    A megawatt-class rack is not an incremental engineering change. It is a different building, a different substation and a different cooling system.

    That is the same arithmetic behind deals like the $45 billion Anthropic–Nscale contract for 460 megawatts in West Virginia. Capacity is being bought years ahead of need.

    What do Nvidia’s numbers say about the thesis?

    They validate it, loudly. Nvidia reported second-quarter fiscal 2027 revenue of $96.2 billion on August 26, up 106% year over year, with data center revenue of $89.0 billion, up 117%, according to the company’s earnings release.

    GAAP gross margin came in at 75.0%. GAAP diluted earnings per share were $2.46.

    Guidance for the current quarter is $108.0 billion, plus or minus 2% — implying another double-digit sequential step up.

    Chief executive Jensen Huang framed it as a regime change: “AI has reached its inflection point. It’s doing useful work. Its tokens are productive and profitable. Now, compute is revenue.”

    One number in that release deserves attention from anyone considering the a16z thesis: guided gross margin slips from 75.0% to 74.0%. Even the company with the most pricing power in the industry is absorbing input costs — a pressure we covered when Nvidia raised AI server prices roughly 15%, with memory the culprit.

    How does this compare to other AI funds?

    It is small in dollars and specific in focus. Where rivals raised general AI megafunds, a16z carved out a thematic slice. SiliconANGLE notes Kleiner Perkins raised $3.5 billion in March 2026 and Thrive Capital raised $10 billion for AI investments.

    Firm Vehicle Size Focus
    Andreessen Horowitz Machine Age Fund $1.1B AI hardware and physical infrastructure
    Kleiner Perkins 2026 vehicle $3.5B General venture, AI-weighted
    Thrive Capital AI vehicle $10B AI, largely late-stage models and apps

    The gap is deliberate. Hardware rounds are capital hungry but the winners are fewer, so a concentrated $1.1 billion can still buy meaningful ownership.

    Dealroom estimates semiconductor and autonomous-machine startups raised roughly $100 billion over the past year. Against that, a16z’s fund is about 1% of the category’s annual intake.

    Why this matters

    Venture capital is a leading indicator of where founders will spend the next five years. When the largest firm in the business stands up a dedicated hardware vehicle, it signals that the software layer looks crowded and the physical layer looks underserved.

    PitchBook analyst Nick Rescigno made the point directly: “Dedicated hardware and robotics funds have existed for years, but when one of the largest firms in venture stands up a fund specifically for that, you pay attention.”

    There is a defensive logic too. PitchBook senior analyst Kaidi Gao noted that “new LLM features could wipe out certain application software AI companies overnight,” pushing investors toward hardware as a hedge.

    For public-market investors, the read-through is that the buildout has more runway than the model-training narrative alone implies. Power, memory and cooling suppliers sit upstream of everything — the same logic behind Broadcom’s up-to-$100 billion debt facility to fund Anthropic chips and the doubling of Etched’s valuation to $21 billion in under a month.

    This post is reporting and analysis, not financial advice.

    What could go wrong with this bet?

    Hardware is a worse venture asset class than software, and nothing in the announcement changes that. Capital intensity is high, build cycles run years, and gross margins outside of Nvidia’s position are thin.

    A $1.1 billion fund also cannot lead many rounds at the scale a16z has been writing. Two $500 million checks in a single month would consume most of it.

    That implies either far smaller positions, heavy syndication, or co-investment from a16z’s larger pools — which makes the headline number more of a branding exercise than a balance-sheet event.

    There is also concentration risk in the thesis itself. Rack-density forecasts assume demand keeps compounding; Nvidia already trimmed one large infrastructure commitment when it cut its OpenAI data center guarantee from $250 billion to $120 billion. Physical assets cannot be repriced overnight the way a SaaS contract can.

    Frequently asked questions

    How big is the a16z Machine Age Fund?

    $1.1 billion, announced August 28, 2026. It covers both early and growth stage investments.

    What does the fund invest in?

    Chips, memory, networking, storage, data centers, power generation and electrical infrastructure, cooling, materials, real estate, robotics and edge AI hardware.

    Who manages the Machine Age Fund?

    General partners Ben Horowitz, Martin Casado, Raghu Raghuram, David Ulevitch and David George, per a16z. Guido Appenzeller, previously CTO of Intel’s data center business, is also on the team.

    Which companies has a16z already backed in this category?

    Named holdings include Unconventional AI, Nexthop, Volta, Atoms, Mind Robotics and Heron Power, alongside older positions in Skydio, Anduril and Waymo.

    Why does rack power consumption matter to investors?

    a16z says racks have gone from 5–10 kW to 100–250 kW and may reach 1 megawatt within three years. That forces new spending on transformers, cooling and buildings — the suppliers the fund targets.

    How does this relate to Nvidia’s latest earnings?

    Nvidia posted $96.2 billion in revenue for the quarter ended August 2026, with data center revenue up 117% year over year. That demand is what the a16z fund is trying to supply.

    Is a16z abandoning software investing?

    No. The Machine Age Fund is a dedicated vehicle alongside the firm’s existing funds. Hardware is more than 20% of deal flow, not all of it.

    The bottom line

    The a16z Machine Age Fund is a $1.1 billion vote that the constraint on AI has moved from algorithms to atoms. The supporting numbers are strong: Nvidia’s $89.0 billion data center quarter, rack power heading toward a megawatt, gigawatt-scale campuses under construction.

    The skepticism is equally simple. A billion dollars does not go far in a category where a16z itself wrote two $500 million checks in one month, and hardware punishes investors who are early.

    Watch two things next: whether other top-tier firms follow with dedicated hardware vehicles, and whether Nvidia’s guided margin compression at 74.0% spreads down the supply chain. If it does, the a16z bet gets more interesting, not less — margin pressure at the top is where component suppliers make their money.

    Sources

  • GLM-5.3-Flash vs Qwen3.8-Flash-Next: Which Cheap Coder Wins

    GLM-5.3-Flash beats Qwen3.8-Flash-Next on coding and costs less to rent: $0.075 per million input tokens on promo through September 9, versus roughly $0.15 on Qwen’s hosted Flash tier. Qwen wins on agentic work and activates 6B parameters against GLM’s 18B, so it is the better model to own. Rent GLM. Self-host Qwen.

    Two Chinese labs shipped a frontier-adjacent open-weight model on the same day. August 26, 2026: Z.ai released GLM-5.3-Flash, Alibaba released Qwen3.8-Flash-Next. Same week, same price bracket, and — as we will get to — very nearly the same architecture.

    This is the comparison that matters for anyone paying an API bill this quarter. Here is what the numbers actually say.

    What are GLM-5.3-Flash and Qwen3.8-Flash-Next?

    Both are sparse mixture-of-experts models built for cheap, long-context, agentic work. GLM-5.3-Flash is the multimodal one with a 1M-token window and an MIT license. Qwen3.8-Flash-Next is a preview of the Qwen4 architecture that activates only 6B parameters per token. Both released August 26, 2026.

    GLM-5.3-Flash: 320B total, 18B active, MIT

    Z.ai’s model carries 320B total parameters with 18B active per token, across 45 layers — 34 linear, 11 full attention — trained on a 30T-token corpus, per LLM-Stats’ launch breakdown.

    It handles text, image and video. The context window is 1,048,576 input tokens with 131,072 output tokens. The license is MIT — the most permissive terms of any model at this capability tier.

    Z.ai claims a 3x attention compute reduction and 4.4x KV cache savings versus the full GLM-5.3.

    Qwen3.8-Flash-Next: 6B active, Qwen4 preview

    Alibaba’s model is 125B in the main body plus a 51B n-gram table and a 4B multi-token-prediction head — roughly 180B stored — but only 6B parameters fire per token. It needs one-ninth the compute of Qwen3.7-Plus.

    Native context is 262,144 tokens, extensible to 1M with YaRN. The license is Qwen Community 1.0, not MIT. Alibaba reports up to 7.6x prefill and 4.9x decoding speedups at 1M tokens.

    How much does each model cost per million tokens?

    GLM-5.3-Flash is cheaper to rent, and it is not close during the promo. Z.ai lists $0.15 input, $0.03 cached input, $0.50 output per million tokens, with 50% off through September 9, 2026. Qwen3.8-Flash-Next has no first-party hosted list price at all — it shipped as weights only.

    Spec GLM-5.3-Flash Qwen3.8-Flash-Next
    Released Aug 26, 2026 Aug 26, 2026
    Total / active params 320B / 18B ~180B stored / 6B
    Native context 1,048,576 tokens 262,144 (1M via YaRN)
    Modalities Text, image, video Text, vision
    License MIT Qwen Community 1.0
    List input / output $0.15 / $0.50 No first-party rate
    Promo input / output $0.075 / $0.25 (to Sep 9)
    Cached input $0.03 Not published
    Nearest hosted sibling Qwen3.8-Flash: $0.16 / $0.47

    What that gap costs on a real workload

    Run 100M input and 20M output tokens a month — a mid-sized coding agent deployment. On GLM’s promo rate that is $7.50 plus $5.00, or $12.50. On Qwen3.8-Flash’s QwenCloud rate of $0.16 / $0.47, it is $16.00 plus $9.40, or $25.40.

    Roughly 2x. Both are rounding errors next to Qwen3.8-Max at $2.00 / $6.00 input-output — the Flash tier undercuts it by more than 10x on input, as DataCamp documented.

    Caching is where GLM pulls further ahead. At $0.03 per million cached input tokens, a repeated 500K-token repo context costs 1.5 cents to re-read. We covered the same dynamic in our breakdown of the cheapest 1M context model.

    Which is better for coding, GLM-5.3-Flash or Qwen3.8-Flash-Next?

    GLM-5.3-Flash wins coding. It scores 63.4 on DeepSWE 1.1 against Qwen’s 58.7, and BenchLM ranks it #13 of 146 on coding versus Qwen at #28. Qwen wins the agentic category decisively — #7 of 140 against GLM’s #36 — so the answer depends on whether your job is writing code or running tools.

    The benchmark split

    Benchmark GLM-5.3-Flash Qwen3.8-Flash-Next
    DeepSWE 1.1 63.4 58.7
    SWE-bench Pro Not published 62.5
    Toolathlon 78.4 Behind GLM
    Terminal Bench 2.1 84.3 Not published
    AutomationBench 48.8 Not published
    GPQA Diamond Not published 91.7
    LiveCodeBench v6 Not published 91.9
    CharXiv-R 89.4% 90.6%
    Agents’ Last Exam Behind Qwen 24.3 pass@1
    BenchLM overall 61.3 (#58/228) 61.3 (#57/228)
    LLM-Stats score 51.6 (#11) 50.5 (#14)

    Note the dead heat at the top line: both land on 61.3/100 at BenchLM, one rank apart. The aggregate hides the split underneath it.

    Where Qwen actually wins

    Qwen’s agentic numbers are the story. SWE-bench Pro 62.5 against Claude Opus 4.6 Max’s 53.4. CoWorkBench 73.9 against 68.2. JobBench 55.7 against 36.6 — a 19-point gap over a frontier closed model.

    It also takes instruction following, ranking #10 of 42 at 91.2. GLM is not measured on that axis.

    Reasoning is Qwen’s weak spot: 35.9 on Humanity’s Last Exam versus Opus’s 40.0. GLM wins the head-to-head on HLE and NL2Repo, per LLM-Stats’ comparison page.

    Why did two rival labs ship the same architecture?

    Because the efficiency math has one answer right now. MarkTechPost’s teardown found both models independently adopted four identical design choices — and the convergence is the real news, more than either model’s scorecard.

    The four shared choices

    • 3:1 linear attention ratio. Three cheap linear layers per full attention layer, compressing history into fixed recurrent states.
    • 4x context compression with sparse attention capped at exactly 2,048 tokens via a learned indexer. Both picked the same number.
    • Four gated residual streams replacing the single-stream transformer, controlled by data-dependent gates.
    • Muon optimizer, with fused matrices split before orthogonalization during training.

    Where they split: RoPE versus NoPE

    GLM dropped rotary position embeddings entirely, relying on linear layers for implicit position. Qwen kept RoPE — after finding that NoPE models “often failed to stop generating” during post-training alignment.

    That is a genuinely useful negative result: a failure mode invisible in pre-training metrics. Not everyone is convinced either way. MiniMax’s ablations found linear attention harms multi-hop reasoning, and M3 uses sparse softmax only.

    Is self-hosting cheaper than the API?

    For GLM-5.3-Flash, no. The FP8 checkpoint is about 306 GiB of weights needing roughly 386 GiB of VRAM — a minimum 8-GPU Hopper node. Two H200s at 282 GiB combined do not fit. For Qwen3.8-Flash-Next at 6B active, the answer flips: fewer active parameters means far cheaper serving at scale.

    The GLM hardware ladder, per LumaDock’s deployment guide:

    • BF16: ~772 GiB VRAM. Multi-node territory.
    • FP8: ~306 GiB weights, ~386 GiB recommended. 8-GPU Hopper node.
    • 4-bit GGUF: ~160 GB before overhead. Two to four GPUs, with quality trade-offs.
    • 1-bit to 3-bit (Unsloth): 100–128 GB combined RAM/VRAM. Mac Studio or DGX class.

    LumaDock’s verdict is blunt: an 8-GPU node costs more per day than most teams spend on the API per month. At $12.50 a month for our example workload, that is not a close call.

    Qwen is the opposite trade. Six billion active parameters is what makes it cheap to serve — the same argument we ran through on Qwen3.8-Max open weights versus API. The catch is the license: Qwen Community 1.0, not MIT. Read it before you build a product on it.

    Which model should you pick?

    Pick by workload, not by leaderboard. GLM for code generation, multimodal input and anything cost-sensitive you plan to rent. Qwen for tool-calling agents, instruction-heavy pipelines and any deployment you intend to own outright.

    Use case Pick Why
    Code generation / repo refactors GLM-5.3-Flash DeepSWE 63.4 vs 58.7; coding #13 vs #28
    Tool-calling agents Qwen3.8-Flash-Next Agentic #7/140; JobBench 55.7
    Document / video understanding GLM-5.3-Flash Multimodal #7/35; video support
    Long-context RAG GLM-5.3-Flash 1M native; $0.03 cached input
    On-prem / air-gapped Qwen3.8-Flash-Next 6B active; runs on far less iron
    Commercial product, license risk GLM-5.3-Flash MIT beats Qwen Community 1.0
    Lowest cost per token today GLM-5.3-Flash $0.075 / $0.25 through Sep 9
    Structured output pipelines Qwen3.8-Flash-Next Instruction following #10/42, 91.2

    Frequently asked questions

    Is GLM-5.3-Flash really MIT licensed?

    Yes. Z.ai released it under MIT, which permits commercial use, modification and redistribution without a revenue threshold. Qwen3.8-Flash-Next ships under Qwen Community License 1.0, which carries its own conditions.

    When does the GLM-5.3-Flash promo price end?

    September 9, 2026. After that, input goes from $0.075 to $0.15 and output from $0.25 to $0.50 per million tokens — a doubling. Budget for it now.

    Can I get Qwen3.8-Flash-Next through an API?

    Not at a first-party list price. It shipped as open weights on Hugging Face. The nearest hosted option is Qwen3.8-Flash on QwenCloud at $0.16 input / $0.47 output; OpenRouter lists a comparable Flash tier at $0.15 / $0.47.

    Which has the bigger context window?

    GLM-5.3-Flash, at 1,048,576 tokens native. Qwen3.8-Flash-Next is 262,144 native and reaches 1M only with YaRN extension.

    Is either one faster?

    BenchLM clocks Qwen3.8-Flash-Next at 73 tokens per second with 30.22s first-token latency; GLM is listed as not measured. Alibaba separately reports up to 7.6x prefill and 4.9x decoding speedups at 1M tokens.

    How do these compare to Western frontier models?

    On agentic coding, favorably. Qwen’s SWE-bench Pro 62.5 beats Claude Opus 4.6 Max’s 53.4. On broad reasoning they still trail — Qwen’s HLE 35.9 versus Opus’s 40.0.

    What hardware do I need to run GLM-5.3-Flash locally?

    An 8-GPU Hopper node for FP8. Community 1-bit to 3-bit GGUF builds run on 100–128 GB of combined RAM/VRAM, with real quality loss.

    The bottom line

    Rent GLM-5.3-Flash. Own Qwen3.8-Flash-Next.

    If you are buying tokens, GLM wins on price, coding accuracy, context length and license, and the $0.03 cached-input rate makes long-context agents genuinely cheap. Move before September 9 and lock in your usage patterns while the promo lasts.

    If you are standing up your own inference, Qwen’s 6B active footprint is the decisive number. Eighteen billion active parameters is three times the serving cost per token, and at scale that swamps a leaderboard gap of five DeepSWE points.

    The one scenario where you should not pick either: a commercial product where license terms carry legal weight and you cannot accept Qwen Community 1.0. There, GLM’s MIT license ends the argument by itself. For a wider look at coding agents, see our comparison of Claude Code vs Codex CLI and GLM-5.3 vs DeepSeek V4 Pro.

    Sources

  • Anthropic Nscale Deal: $45 Billion for 460 Megawatts in West Virginia

    The Anthropic Nscale deal commits Anthropic to pay Nscale roughly $45 billion over six years for about 460 megawatts of computing capacity at the Monarch Compute Campus in Mason County, West Virginia. CNBC and Bloomberg reported the agreement on August 26, 2026. The first building comes online in late 2027 on Nvidia Vera Rubin hardware, weeks before both companies hope to go public.

    Two companies that have never turned in a public quarterly report just signed one of the largest private compute contracts on record. Both are pre-IPO. The timing is not an accident.

    How big is the Anthropic Nscale deal?

    The Anthropic Nscale deal is worth approximately $45 billion over six years, covering about 460 megawatts of leased capacity, according to CNBC, which cited people familiar with the matter. Bloomberg reported the same figure. That works out to roughly $7.5 billion a year, or about $16.3 million per megawatt over the contract life.

    West Virginia MetroNews reported on August 27 that Anthropic will take the entire first building at the campus, roughly one-third of the three-building site.

    Deal terms at a glance

    Term Detail Source
    Contract value ~$45 billion CNBC, Bloomberg
    Duration 6 years CNBC
    Capacity leased ~460 MW CNBC
    Site Monarch Compute Campus, Mason County, WV WV MetroNews
    Hardware Nvidia Vera Rubin NVL72 Reported specs
    Online date Late 2027 CNBC
    Nscale IPO target September 2026 Bloomberg
    Anthropic IPO target As early as October 2026 Reported filings

    Where is the Monarch Compute Campus?

    Monarch sits north of Point Pleasant in Mason County, West Virginia, and is being developed by Nscale. Three data center buildings are planned. Anthropic’s 460 megawatts fill the first one. The campus runs on a natural gas microgrid rather than grid power or hydro.

    Governor Patrick Morrisey called the reports “an extraordinary vote of confidence in our state and the strategy we put in place,” according to WV MetroNews.

    What West Virginia actually gets

    An Ernst & Young study released the same week put numbers on the local impact. They are meaningful for a county of fewer than 30,000 people, and modest against a $45 billion contract.

    • 6,000 jobs over two years, mostly construction
    • 700 permanent positions in Phase 1
    • $105 million in projected tax revenue
    • $87 million of that to Mason County schools, $13 million to county government, $5 million to EMS and public safety
    • $500,000 donated by Nscale to the Mason County Board of Education for workforce training

    That is roughly $150,000 of projected public revenue per permanent job, spread across the life of the campus. Environmental advocates have raised air quality objections to the gas microgrid.

    Why is Anthropic spending this much on compute?

    Because demand outran its own forecasts. Anthropic’s Q2 2026 revenue reached about $11.6 billion, CNBC reported on August 15, more than doubling the prior quarter and passing OpenAI’s quarterly revenue for the first time. Compute is the input constraint on that curve.

    The Nscale contract is the latest layer on a stack that has been building for nearly a year.

    Anthropic’s compute commitments

    1. Fluidstack — $50 billion, announced November 2025, Texas and New York sites, per Anthropic’s own newsroom
    2. SpaceX — reported $1.25 billion per month, May 2026, the only deal that delivered immediate capacity
    3. AMD — $5 billion, July 2026
    4. Volta — $10 billion, early August 2026, Norway
    5. Nscale — $45 billion, August 2026, West Virginia

    Add the expanded Amazon, Google and Broadcom arrangements from April 2026 and Anthropic has committed well over $110 billion in disclosed compute obligations against a business that has only just posted its first positive operating quarter. We covered the debt side of that build in the Broadcom financing package for Anthropic chips.

    Here is the uncomfortable arithmetic. Anthropic’s entire Q2 revenue would cover about fifteen months of the Nscale contract alone. Every other commitment sits on top of that.

    What does the deal mean for Nscale’s IPO?

    It anchors it. Nscale, a UK infrastructure company founded in 2024, told prospective investors it held about $51 billion in contracted revenue, Bloomberg reported on August 6. It is targeting a US listing as soon as September with Goldman Sachs and JPMorgan advising. PYMNTS reported the raise could reach $3 billion.

    Nscale’s actual recognized revenue is a different order of magnitude: roughly $33 million for all of 2025, about $37 million in Q1 2026, and more than $100 million in Q2 2026.

    The concentration problem nobody has solved

    If the $45 billion sits inside that $51 billion backlog, Anthropic is roughly 88% of Nscale’s entire book. If it does not, the backlog nearly doubles overnight and Anthropic is still just under half. Either way, one customer defines the company.

    Then there is the deal that did not happen. Microsoft signed a 1.35-gigawatt letter of intent at the same Monarch campus in March 2026 and never converted it into a binding contract. Microsoft has not publicly explained why.

    Anthropic’s agreement is reportedly binding, which is the material difference. But an IPO prospectus that leans on a single six-year contract, at a site a hyperscaler walked away from, on chips that have not yet shipped at commercial scale, is a specific kind of bet.

    Why this matters

    The AI infrastructure market has moved from buying compute to underwriting it. Nscale is not selling capacity it owns. It is selling capacity it will build using the contract as collateral. Anthropic’s signature is the balance sheet.

    That is the same structure showing up across the sector, and the terms keep getting revised. Nvidia trimmed its own exposure earlier this month, as we noted when it cut its OpenAI data center guarantee from $250 billion to $120 billion. Hardware costs are moving too — see the 15% Nvidia server price increase.

    For investors, three things follow. Neoclouds are becoming credit instruments rather than cloud businesses. Customer concentration is the risk factor that actually matters in this cohort, not utilization. And a 2027 delivery date means the first real test of these contracts is still more than a year out.

    Anthropic’s own listing, which we examined when its valuation reached $2 trillion, will have to explain these obligations in a prospectus. That document will be more informative than any deal headline.

    This post is reporting and analysis, not financial advice.

    Frequently asked questions

    How much is the Anthropic Nscale deal worth?

    About $45 billion over six years, according to CNBC and Bloomberg, both citing people familiar with the agreement. Neither company has published the contract.

    How much power does Anthropic get?

    Roughly 460 megawatts, filling the first of three planned buildings at the Monarch Compute Campus in Mason County, West Virginia.

    When does the capacity come online?

    Late 2027. The site will run Nvidia Vera Rubin NVL72 systems, which have not yet shipped at commercial scale.

    What is Nscale?

    A UK-based AI infrastructure company founded in 2024. It reports about 831 megawatts of active and contracted power and roughly 25,000 active GPUs, mostly Nvidia Blackwell.

    Why did Microsoft walk away from the same site?

    Microsoft signed a 1.35-gigawatt letter of intent at Monarch in March 2026 and did not convert it to a binding agreement. It has given no public explanation.

    Is Nscale profitable?

    No published figure suggests so. Recognized revenue was roughly $33 million in 2025 and more than $100 million in Q2 2026, against a contracted backlog of about $51 billion.

    When are the IPOs?

    Nscale is targeting September 2026. Anthropic filed confidentially in June 2026 and is reported to be targeting a Nasdaq listing as early as October 2026.

    The bottom line

    The Anthropic Nscale deal is a $45 billion vote of confidence placed by one pre-IPO company in another, five weeks before the first of them tries to list. It gives Nscale the anchor tenant its prospectus needs and gives Anthropic capacity it will not touch until late 2027.

    Watch three things. Whether Nscale’s September filing discloses the concentration honestly. Whether Vera Rubin ships on schedule. And whether Anthropic’s October prospectus reconciles more than $110 billion in compute obligations against $11.6 billion of quarterly revenue.

    The contracts are signed. The capacity is not built.

    Sources

  • Model Hardware Standard: Anthropic Cuts Lab Setup to 8 Hours

    Anthropic released the Model Hardware Standard on August 27, 2026 — a research preview that lets Claude and rival models drive lab robots, pipettes and factory arms through a single spec. Early testers cut integration from weeks to hours. Carnegie Mellon stood up a serial dilution workflow in 8 hours. QuEra took a laser recovery routine from 58% success to 99.3%. No pricing, no revenue, no open-source date.

    Anthropic has spent two years selling tokens that move text. This one moves matter.

    The company published the Model Hardware Standard, or MHS, as a research preview on August 27. It is a specification, not a product — closer to a plug shape than to a machine. And that is exactly the point.

    What is the Model Hardware Standard?

    The Model Hardware Standard is a shared specification that tells an AI agent what a physical device can do and, more importantly, what it must never do. Vendors ship a driver. The agent reads and writes through simple primitives. Anthropic is running it as an invitation-only research preview.

    Per Anthropic’s own announcement, MHS works with any device that exposes a programmable interface. It is model-agnostic by design — Claude is not required.

    That last detail matters more than the demos. Anthropic is not shipping a robot. It is trying to own the socket every robot plugs into.

    How MHS actually works

    A vendor writes one standardized driver. That driver publishes device discovery in a common format and exposes controls, sensor values and safety limits through a shared memory dictionary.

    The agent then reaches the hardware through one of three paths: the Model Context Protocol, a command line interface, or generated code files. Same device, three levels of abstraction.

    Safety limits live in the driver, not in the prompt. Anthropic gives the example of blocking excess laser power at the device layer — so a confused model cannot talk its way past a hardware ceiling.

    Where MCP ends and the Model Hardware Standard begins

    MCP, which Anthropic debuted in 2024, connects models to software: databases, ticket systems, file stores. If you have followed our coverage of how Agent Skills and MCP split the token bill, the architecture will look familiar.

    MHS extends the same logic to things with motors. Anthropic technical staff member Alek Kemeny put it bluntly to TNW: “What MCP did for software, MHS will do for the hardware world.”

    Kemeny has described MCP elsewhere as “kind of like the USB for AI to software connection.” MHS is the industrial-grade version of that pitch.

    What did the early tests actually prove?

    Six organizations ran MHS against real equipment before launch, and the reported results are specific rather than vague. The headline claim is time: Anthropic says MHS “reduces this integration work to hours or minutes,” against a baseline Genentech described as weeks or months of manual work.

    The most concrete number came from quantum computing. QuEra used MHS to rebuild a laser stabilization routine, moving from 58% success at 150 seconds to 99.3% success at 6 seconds — a 25x speedup on a task that was already mostly failing.

    Organization What was automated Reported result
    QuEra Computing Laser stabilization recovery 58% to 99.3% success; 150s to 6s
    Carnegie Mellon Serial dilution workflow 3x faster; 8-hour setup vs. weeks
    Tetsuwan Scientific qPCR liquid handling 9,143 dispenses across 300 transfer types
    Genentech BCA protein assay tuning Converged at ~140 µL/s (water), 10 µL/s (BSA)
    University of Washington Multi-instrument bench Six instruments connected in under a week
    HHMI Janelia Co-development partner Reference implementation

    Anthropic also says it tested six failure conditions on purpose: missing plate, rotated plate, reader busy, disconnected camera, unreachable device, emergency stop. That is a short list for anything touching a factory floor.

    The number that should give buyers pause

    Every figure above comes from partners Anthropic selected and published. None of it is independently benchmarked, and there is no public failure rate across the full preview cohort.

    A 99.3% success rate on a laser is excellent in a lab. On a production line running 20,000 cycles a shift, it is 140 faults.

    Who is backing the Model Hardware Standard?

    Anthropic named ten hardware vendors and six research institutions at launch. The vendor list is the commercially interesting half, because those are the companies that would have to ship MHS drivers in firmware for the standard to matter.

    Vendors listed by Anthropic as supporting or planning support:

    • Amazon Web Services (Strands Robots library)
    • Universal Robots
    • Doosan Robotics
    • Danaher
    • QIAGEN
    • Tecan
    • Automata
    • MBF Bioscience
    • Hugging Face (LeRobot)
    • Raspberry Pi

    Research users include Genentech, Carnegie Mellon, the University of Washington’s Baker and Pinglay labs, HHMI Janelia, QuEra and Tetsuwan Scientific.

    Jonah Cool, Anthropic’s head of partnerships and deployment of science, told Fortune that lab equipment “suffers from proprietary solutions that are very brittle,” and that the goal is to “avoid vendor lock-in for scientists.”

    Read that again from a vendor’s chair. Anthropic is asking Danaher, QIAGEN and Tecan to help dismantle the integration moat that protects their service revenue.

    How much does the Model Hardware Standard cost?

    Nothing, for now — and that is the strategy. MHS is free during the research preview, gated by an invitation waitlist at modelhardwarestandard.com. Anthropic says it will open-source the framework after the preview, but has published no date, no license and no commercial terms.

    Standards are loss leaders. The money is downstream, in the tokens burned by agents that run instruments around the clock.

    An overnight experiment is a 12-hour inference session. Multiply that by a few thousand labs and the economics start to look like a metered utility rather than a chat subscription.

    Who wins and who loses financially?

    The winners are frontier labs with agent products and the robotics vendors with thin software teams. The losers are instrument makers whose margins depend on proprietary integration, and the systems integrators paid by the week to wire benches together.

    Winners

    Anthropic first. The company was reported at a $2 trillion valuation earlier this month, and a hardware standard extends its distribution into a market where it currently sells nothing.

    Robot arm vendors win cheaply. Universal Robots and Doosan get an agent interface without building an AI stack — the same trade that made Unitree’s IPO pop 629% a bet on hardware plus somebody else’s brains.

    Cloud providers win the runtime. AWS shipped Strands Robots support on day one for a reason.

    Losers

    Integration consultancies are the clearest casualty. If a Carnegie Mellon bench goes from several weeks to 8 hours, that is billable work evaporating.

    Proprietary lab software is next. MarketsandMarkets valued lab automation at $6.60 billion in 2026, growing to $8.62 billion by 2031 at a 6.6% CAGR — a slow market where vendors defend share through lock-in, not growth.

    A commoditized driver layer is precisely the thing that breaks that defense.

    Is the Model Hardware Standard safe enough to run a factory?

    Not yet, and Anthropic says so. The company acknowledged that large language models “still lack physical intuition,” and states that safety evaluations are being built during the preview rather than before it. Human approval workflows exist for high-risk actions, but the physical safety roadmap is unfinished.

    The Register, which covered the launch on August 28, raised the obvious dual-use question: a universal spec for driving instruments does not care what the instrument is for.

    The January 2027 regulatory deadline

    EU Machinery Regulation 2023/1230 takes effect on January 20, 2027. It is the first EU rule to cover AI-based safety functions and self-evolving machine behavior.

    TNW notes the awkward implication: an MHS file that constrains how a machine may operate could itself qualify as a regulated safety component. That would put liability on whoever wrote the driver.

    Anthropic has not said who that is. Five months out from the deadline, this is the unpriced risk in the whole announcement.

    How does this fit Anthropic’s broader agent push?

    MHS is the physical endpoint of a strategy that has been visible all year in software. Anthropic has been widening what an agent can touch, from computer-use agents driving desktops to skills that compress tool definitions.

    The competitive timing is not subtle either. Fortune reported that Hugging Face shipped a robotic duck the same day, and that Nvidia is pursuing a $13 billion acquisition of the company.

    Physical AI is where the capital is rotating. German humanoid maker NEURA Robotics raised up to $1.4 billion in Series C funding this year, per TNW.

    Frequently asked questions

    Is the Model Hardware Standard open source?

    Not yet. Anthropic says it intends to open-source the framework after the research preview, but has published no date or license. Drivers built during the preview are being made available for reuse.

    Does MHS only work with Claude?

    No. Anthropic describes MHS as model-agnostic, meaning OpenAI models and open-weight models can drive MHS devices. Whether rival labs adopt a spec authored by a competitor is a separate question.

    How is MHS different from MCP?

    MCP connects models to software. MHS connects them to physical devices, and adds device-level safety limits, sensor state and discovery. MCP is one of three ways to reach an MHS device, alongside a CLI and generated code.

    Can I use it today?

    Only by invitation. Access runs through a waitlist at modelhardwarestandard.com, and Anthropic has described early access as a “handful” of labs and manufacturers in biotech, robotics and quantum computing.

    What hardware is supported?

    Anything with a programmable interface, in principle. In practice, ten named vendors — including Universal Robots, Danaher, QIAGEN, Tecan and Raspberry Pi — are supporting or planning support. Older instruments without a programmable interface are out of scope.

    What is the biggest risk?

    Regulation and liability. EU Machinery Regulation 2023/1230 applies from January 20, 2027, and MHS constraint files may count as regulated safety components — with no clarity yet on who carries responsibility when an agent-driven machine injures someone.

    The bottom line

    The Model Hardware Standard is the most strategically aggressive thing Anthropic has shipped this year, and it contains no product.

    The engineering claims are credible and unusually specific. A 58% to 99.3% jump on QuEra’s laser routine and an 8-hour Carnegie Mellon integration are not marketing numbers. They are the kind of figures a skeptical buyer can go test.

    But every one of them came from a partner Anthropic chose. There is no pricing, no open-source date, no independent benchmark and no answer on who is liable when a driver written by a language model moves a robot arm into a person.

    The verdict: treat MHS as a distribution land-grab, not a revenue event. If ten vendors becomes fifty by January, Anthropic will own the plug shape for physical AI and collect inference rent on every machine that uses it. If the EU deadline arrives with the liability question still open, the same vendors will quietly wait it out.

    Watch the driver count, not the demos.

    Sources

  • Nvidia Hugging Face Acquisition: $12.9 Billion at 80x Revenue

    The Nvidia Hugging Face acquisition values the open-source model hub at $12.9 billion, according to The Information — roughly 80 times its ~$150 million in annualized revenue. Hugging Face turned down a $500 million Nvidia investment at a $7 billion valuation less than a year ago. Neither company has confirmed the deal. It would be Nvidia’s second-largest acquisition ever, behind the $20 billion Groq purchase.

    How much is Nvidia paying for Hugging Face?

    Nvidia has agreed to pay approximately $12.9 billion for Hugging Face, The Information reported on August 27, citing a person familiar with the transaction. The deal is agreed but not signed. It can still collapse.

    CNBC and Fortune both matched the report the same day. Business Insider first reported Nvidia’s takeover interest.

    Neither Nvidia nor Hugging Face responded to requests for comment, per Quartz. That silence matters — nothing here is a signed, disclosed transaction yet.

    What Hugging Face was worth before

    Hugging Face last priced itself at $4.5 billion in a 2023 round led by Salesforce Ventures, with Alphabet’s GV, IBM Ventures and — notably — Nvidia participating, according to TechCrunch.

    In late 2025, Nvidia offered $500 million at a $7 billion valuation. Hugging Face said no. TechCrunch reports the company declined because it did not want a single dominant investor.

    Roughly nine months later, it is selling outright to that same investor for nearly double the valuation it rejected.

    Date Event Valuation
    2023 Series funding led by Salesforce Ventures $4.5 billion
    Late 2025 $500M Nvidia investment offer — declined $7 billion
    Aug 27, 2026 Reported acquisition agreement $12.9 billion

    Why is Nvidia buying an open-source model hub?

    Nvidia is buying distribution, not revenue. Hugging Face is where developers publish, discover and download open models — Tom’s Hardware calls it a “GitHub-like repository” for AI. Owning the shelf is worth more to Nvidia than the $150 million the shelf currently earns.

    The strategic timing is not subtle. Nvidia’s largest customers are building silicon that competes with its own.

    The lock-in play

    Hugging Face’s Inference Endpoints today support AWS Inferentia, AMD Instinct, Google TPU, Intel CPUs and Nvidia accelerators, per Tom’s Hardware. It is deliberately vendor-neutral.

    Under Nvidia, that neutrality is the first thing analysts expect to erode. If the default deployment path for every popular open model points at CUDA, Nvidia defends its installed base at the exact layer where switching decisions get made.

    That threat is real. OpenAI, Google, Amazon and Anthropic are all shipping or funding custom accelerators — see our coverage of OpenAI’s Jalapeño chip and the Broadcom debt package funding Anthropic’s silicon.

    The cloud re-entry play

    Nvidia scaled back DGX Cloud roughly a year ago, TechCrunch notes. Hugging Face gives it a consumer-facing compute surface again — and somewhere to route the capacity Nvidia has committed to but not fully sold.

    That is the least-discussed part of the rationale and possibly the most financially concrete one.

    Is $12.9 billion too much for $150 million in revenue?

    On the numbers, yes — by any conventional standard. Tom’s Hardware puts the deal at roughly 80 times forward revenue. Software acquisitions at 15–20x are already considered rich.

    Hugging Face’s revenue is growing fast. TechCrunch reports it moved from about $100 million to about $150 million annualized in roughly two months, and Tom’s Hardware says paying subscribers doubled in the first half of 2026. CEO Clem Delangue told TechCrunch last month the company was “close to profitability.”

    Here is the skeptical read. Nvidia is paying a strategic premium for neutrality it intends to end. The moment developers believe Hugging Face is a CUDA storefront rather than a Switzerland, some of them leave — and the asset Nvidia bought is worth less than the asset it paid for. AMD, Google and the open-weights community have every incentive to fund an alternative registry.

    Ten-year-old infrastructure businesses with $150 million in revenue do not usually command $12.9 billion. They command it when the buyer is defending a franchise.

    How does this fit Nvidia’s acquisition spree?

    It is the second-largest deal Nvidia has ever done, and the third multi-billion-dollar AI purchase in nine months. Nvidia has stopped behaving like a component supplier and started behaving like a platform consolidator.

    Target Reported price Announced What it buys
    Groq ~$20 billion Dec 2025 Inference architecture (LPU)
    Hugging Face $12.9 billion Aug 2026 Open-model distribution
    Poolside ~$6 billion Aug 2026 Model training capability

    The Groq deal — about $20 billion, reported by CNBC in December 2025 — was Nvidia’s largest on record. We covered the $6 billion Poolside purchase earlier this month.

    Nvidia can afford all of it in cash. Its Q2 fiscal 2027 results, for the quarter ended July 26, 2026, show $96.2 billion in revenue, up 106% year over year, and $59.7 billion in GAAP net income. Cash, marketable debt and marketable equity securities totaled roughly $99.3 billion.

    Why this matters

    Three things follow from this deal, and none of them are about Hugging Face.

    • The competitive threat is now priced. Nvidia is spending real money to defend against customers building their own chips. That is an admission the threat is material.
    • Open-source AI just got an owner. The default distribution point for open models moves inside a hardware vendor. Expect immediate pressure for a neutral alternative.
    • Strategic multiples are back. Eighty times revenue is a 2021-style number appearing in 2026, funded by operating cash rather than cheap debt.

    For investors, the read is about defensive capital allocation. Nvidia guided to about $108 billion for the current quarter — excluding any China data center compute revenue. A company growing that fast does not spend $12.9 billion on a $150 million business unless it sees a hole in the moat.

    It also fits a pattern of Nvidia paying to control adjacent chokepoints, much as Stripe paid over $7 billion for OpenRouter to sit on the AI token toll road. Note, too, that Nvidia recently cut its OpenAI data center guarantee from $250 billion to $120 billion — capital is being redirected, not simply added.

    This post is reporting and analysis, not financial advice.

    Frequently asked questions

    Is the Nvidia Hugging Face acquisition confirmed?

    No. The Information reported an agreement on August 27, 2026, and CNBC, Fortune and TechCrunch matched it. Neither company has commented publicly, and the deal is not signed.

    How much revenue does Hugging Face generate?

    Roughly $150 million annualized, up from about $100 million two months earlier, according to TechCrunch. The $12.9 billion price is about 80 times that figure.

    Why did Hugging Face reject Nvidia before?

    It declined a $500 million investment at a $7 billion valuation in late 2025 because it did not want a single dominant investor, TechCrunch reported.

    Will Hugging Face still support AMD and Google chips?

    Unknown. Its Inference Endpoints currently support AWS Inferentia, AMD Instinct, Google TPU, Intel CPUs and Nvidia accelerators. Nvidia has not said whether that continues.

    Is this Nvidia’s biggest acquisition?

    No. The roughly $20 billion Groq deal announced in December 2025 remains its largest, per CNBC. Hugging Face would rank second.

    Could regulators block the deal?

    No formal review has been reported. Antitrust scrutiny is plausible given Nvidia’s accelerator share and the platform’s role in model distribution, but nothing has been filed publicly.

    Can Nvidia pay cash?

    Comfortably. It reported $59.7 billion in GAAP net income in a single quarter and about $99.3 billion in cash and marketable securities as of July 26, 2026.

    The bottom line

    Nvidia is paying roughly 80 times revenue to own the front door of open-source AI. The financial case is thin; the defensive case is obvious.

    Watch three things next: whether the deal is actually signed, whether Hugging Face keeps supporting rival accelerators, and how quickly a neutral competitor gets funded. The first tells you if this is real. The second and third tell you whether $12.9 billion bought a moat or a melting asset.

    Sources

  • Cheapest 1M Context Model: GLM-5.3-Flash vs Gemini 3.7 Flash

    GLM-5.3-Flash is the cheapest 1M context model worth running in production. Z.ai lists it at $0.15 per million input tokens against $0.75 for Gemini 3.7 Flash — five times cheaper — while scoring 57 on the Artificial Analysis Intelligence Index versus Gemini’s 56. Google keeps two real advantages: raw throughput and vision. Everything else favors the open-weights challenger.

    Z.ai shipped GLM-5.3-Flash on August 26, 2026, thirteen days after Google made Gemini 3.7 Flash generally available. Both models advertise a 1,048,576-token context window. Both target agentic coding and long-document work.

    The gap is price. And at 1M-token scale, price is the entire product decision.

    What is GLM-5.3-Flash?

    GLM-5.3-Flash is a natively multimodal mixture-of-experts model with 320 billion total parameters and 18 billion active per token, released under an MIT license. It routes each token through 8 of 288 experts across 45 layers, ships in native FP8, and holds a 1,048,576-token context window.

    That active-parameter count is the whole story. Z.ai is charging flagship-tier context for a model that only lights up 18B weights per forward pass.

    The architecture behind the price

    The model combines KDA linear-attention layers with NoPE sparse MLA layers. Per MarkTechPost’s launch coverage, that combination delivers roughly 3x less attention compute and a 4.4x smaller KV cache than GLM-5.3.

    KV cache is what makes long context expensive to serve. Shrink it 4.4x and you can price a 1M window like a short one.

    The jump over the previous generation is not cosmetic. Z.ai’s own numbers put DeepSWE v1.1 at 63.4%, up from 46.2% on GLM-5.2, and AutomationBench at 48.8%, up from 26.2% — a 22.6-point gain in one release cycle, according to LLM Stats.

    How much does the cheapest 1M context model actually cost?

    GLM-5.3-Flash lists at $0.15 per million input tokens and $0.50 output, with cached input at $0.03. Gemini 3.7 Flash lists at $0.75 input and $3.75 output on Google’s own model page. That is 5x on input and 7.5x on output, before any discount either side is running.

    Spec GLM-5.3-Flash Gemini 3.7 Flash
    Released Aug 26, 2026 Aug 13, 2026
    License MIT open weights Proprietary API
    Parameters 320B total / 18B active Undisclosed
    Input context 1,048,576 tokens 1,048,576 tokens
    Max output 131,072 tokens 65,536 tokens
    Input / 1M $0.15 $0.75
    Output / 1M $0.50 $3.75
    Cached input / 1M $0.03 $0.06 (Vertex)
    AA Intelligence Index 57 56
    Output speed 50.2 tok/s 301 tok/s
    Time to first token 1.47s 3.83s

    Pricing from Z.ai list rates and Google DeepMind’s Gemini Flash page. Speed and index figures from Artificial Analysis and Requesty’s Vertex listing. Resellers differ: OpenRouter lists GLM-5.3-Flash at $0.075 / $0.25 and Gemini 3.7 Flash at $0.375 / $1.875.

    What it costs to fill the window once

    Push a full 1,048,576-token context through each model, one time, and the arithmetic is brutal.

    GLM-5.3-Flash: $0.157. Gemini 3.7 Flash: $0.786. Same window, same task, a $0.63 difference per call.

    Run that 10,000 times a month — a modest document-processing pipeline — and you are looking at $1,573 versus $7,864. The $6,291 monthly delta is a headcount line item, not a rounding error.

    For context on how wide the field has gotten, Morph’s context-window survey clocked a 71x spread between the cheapest and priciest 1M window on the market, from $0.14 on DeepSeek V4 Flash to $10.00 on Claude Fable 5.

    The January 2027 price cliff

    Google’s $0.75 / $3.75 is an introductory rate. Its own page states the promotion expires December 31, 2026, after which Gemini 3.7 Flash reverts to $1.50 per million input and $7.50 per million output.

    On January 1, filling that same 1M window costs $1.57 on Gemini. Against GLM’s $0.157, that is a clean 10x.

    Z.ai is running a promotion too — 50% off through September 9, 2026 — but its post-promo list price is the $0.15 already quoted. One vendor’s discount expires into a doubling. The other’s expires into the number on the page.

    Which is better for coding agents, GLM-5.3-Flash or Gemini 3.7 Flash?

    Gemini 3.7 Flash wins the coding benchmarks by margins too small to justify a 7.5x output bill. It leads Terminal-Bench 2.1 85.8% to 84.3% and DeepSWE v1.1 65.3% to 63.4%. GLM takes HLE 55.3% to 53.6% and destroys Gemini on AutomationBench, 48.8% to 30.4%.

    A 1.5-point Terminal-Bench edge is inside the noise band of most agent harnesses. An 18.4-point AutomationBench gap is not.

    AutomationBench measures multi-step tool use and workflow completion — the thing you actually buy an agent model for. GLM-5.3-Flash scores 60% higher there in relative terms.

    Coding agents also burn output tokens, not input tokens. A long agentic run is thousands of generated tokens per step. That is precisely the axis where Gemini costs 7.5x more.

    Our earlier breakdown of GLM-5.3 against DeepSeek V4 Pro found the same pattern in the open-weight tier: near-parity capability, order-of-magnitude price separation.

    Where Gemini 3.7 Flash still wins

    Google has genuine leads that no discount closes:

    • Throughput: 301 tokens/second median output versus 50.2 for GLM-5.3-Flash — 6x faster generation.
    • Vision: BabyVision 70.9% against GLM’s 53.4%, a 17.5-point gap that Z.ai does not dispute.
    • Long-context recall: GDM-MRCR v2 at 128k scores 97.0%, among the strongest retrieval numbers published this year.
    • Desktop agents: OSWorld-2.0 at 47.9% and Code Arena at 1588 Elo for web development.
    • Vertical accuracy: Harvey LAB-AA at 90.7% on legal reasoning tasks.

    GLM does answer faster on the first token — 1.47s versus 3.83s — which matters for interactive chat. But once generation starts, Gemini pulls away hard.

    Is GLM-5.3-Flash worth it for multimodal work?

    Only for charts and documents, not for general vision. GLM-5.3-Flash posts 78.0% on Chartography against DeepSeek-V4-Flash-Vision-Exp’s 64.3%, but trails Gemini 3.7 Flash badly on BabyVision, 53.4% to 70.9%. Structured visual data is a strength. Open-ended image understanding is not.

    It also scores 62.4% on OfficeQA Pro, which points at the same conclusion: business documents, spreadsheets, slides and charts are where the multimodal stack earns its keep.

    If your pipeline reads invoices, financial statements or dashboards, GLM handles it at a fifth of the price. If it captions arbitrary photos, pay Google.

    We ran similar math on DeepSeek’s vision model against Claude Opus 4.8, where the per-image gap ran 23x. Cheap vision is now a solved category — you just have to match the model to the image type.

    Which model should you buy for your workload?

    Pick on token mix, not on leaderboard position. Input-heavy jobs at 1M scale go to GLM-5.3-Flash on cost alone. Latency-critical streaming and general vision go to Gemini 3.7 Flash. Coding agents are close on quality and lopsided on price.

    Use case Buy Why
    Bulk document / RAG ingestion GLM-5.3-Flash $0.157 vs $0.786 per full 1M window
    Long-horizon coding agents GLM-5.3-Flash AutomationBench 48.8 vs 30.4; 7.5x cheaper output
    Real-time chat / streaming UX Gemini 3.7 Flash 301 tok/s vs 50.2 tok/s
    General image understanding Gemini 3.7 Flash BabyVision 70.9 vs 53.4
    Charts, invoices, office docs GLM-5.3-Flash Chartography 78.0; OfficeQA Pro 62.4
    Needle-in-haystack retrieval Gemini 3.7 Flash GDM-MRCR v2 at 97.0%
    Data that cannot leave your VPC GLM-5.3-Flash MIT weights, self-hostable
    Terminal-Bench maximalists Gemini 3.7 Flash 85.8 vs 84.3 — for a 5x premium

    Should you self-host GLM-5.3-Flash instead?

    Only above roughly 100 million tokens a month. The FP8 checkpoint is 306 GiB of weights before KV cache and needs NVIDIA Hopper or newer. That is a multi-GPU node running continuously against an API bill of $0.15 per million input tokens.

    The MIT license is the real asset here, not the savings. It permits commercial use, modification and redistribution with no revenue thresholds — which is what makes GLM viable for regulated buyers who cannot route customer data through a third-party API.

    Weights are published on Hugging Face as zai-org/GLM-5.3-Flash. Our Qwen3.8-Max self-hosting cost analysis laid out the crossover math in detail; the shape is unchanged, only the weight file got smaller.

    Frequently asked questions

    Is GLM-5.3-Flash actually the cheapest 1M context model?

    Not quite. DeepSeek V4 Flash fills a 1M window for about $0.14 against GLM’s $0.157. But GLM scores 57 on the Artificial Analysis Intelligence Index and adds native multimodality, which makes it the cheapest capable one.

    How much cheaper is GLM-5.3-Flash than Gemini 3.7 Flash?

    Five times cheaper on input ($0.15 vs $0.75 per million) and 7.5 times cheaper on output ($0.50 vs $3.75). After Google’s introductory pricing expires December 31, 2026, the input gap widens to 10x.

    Does GLM-5.3-Flash really have a 1M context window?

    Z.ai specifies 1,048,576 tokens. OpenRouter lists its routed endpoint at 1,310,720 tokens with 131,072 max output — double Gemini 3.7 Flash’s 65,536-token output ceiling.

    Which model is faster?

    Gemini 3.7 Flash generates 6x faster at 301 tokens per second versus 50.2. GLM-5.3-Flash responds faster initially, at 1.47 seconds to first token against Gemini’s 3.83 seconds.

    Are these benchmark scores independently verified?

    Partly. The Artificial Analysis Intelligence Index scores are third-party. The Terminal-Bench, DeepSWE and AutomationBench figures are vendor self-reported on both sides — LLM Stats flags this explicitly for GLM-5.3-Flash.

    Can I use GLM-5.3-Flash commercially?

    Yes. The weights ship under an MIT license, which permits commercial use, modification and redistribution without revenue caps or usage restrictions.

    What happens to Gemini 3.7 Flash pricing in 2027?

    Google’s model page states the introductory rate ends December 31, 2026, moving to $1.50 per million input tokens and $7.50 per million output from January 1, 2027.

    The bottom line

    Buy GLM-5.3-Flash. For any workload dominated by input tokens or agent output tokens, it is the correct default — 5x to 7.5x cheaper at an intelligence index one point above Gemini 3.7 Flash, with a bigger output ceiling and weights you can take in-house.

    Keep Gemini 3.7 Flash for exactly two jobs: user-facing streaming where 301 tokens per second is the product, and general image understanding where 17.5 BabyVision points decide whether the feature works at all.

    The broader signal matters more than either model. Google discounted a Flash-tier model and still got undercut 5x by open weights released thirteen days later. Google’s own price sheet says that gap widens to 10x in four months.

    If you are still routing 1M-token jobs through a proprietary Flash endpoint in 2027, you are paying a tenfold convenience tax. Compare that against our Gemini 3.7 Flash versus Claude Sonnet 5 cost-per-point analysis and the direction is unmistakable.

    Sources

  • OpenAI Jalapeño Chip Beats Blackwell 1.9x Per Watt — Ships 2027

    The OpenAI Jalapeño chip, the company’s first custom inference ASIC, delivered 1.5x to 1.9x more AI work per watt than Nvidia’s Blackwell systems in SemiAnalysis InferenceX tests published August 25, 2026. It draws 700W against GB300’s 1,400W and cut end-to-end latency by up to 3.6x. The catch: these are engineering samples. Volume deployment does not arrive until 2027.

    What is the OpenAI Jalapeño chip?

    The OpenAI Jalapeño chip is a custom inference accelerator co-developed with Broadcom and fabricated on TSMC’s N3P node. It is built to serve tokens, not train models. OpenAI published its first third-party benchmarks this week, and they are better than any first-generation silicon has a right to be.

    The headline spec: 13.4 PFLOPS of MXFP4 compute at a 700W rating, paired with HBM4 running at 15.4 TB/s of bandwidth. In sustained operation the part draws under 550W, according to the benchmark data reported by ForkLog.

    Nvidia’s GB200 rack unit pulls 1,200W. GB300 pulls 1,400W. Rubin sits between 900W and 1,150W. Jalapeño is doing its work in roughly half the power envelope.

    The timeline is the real story

    OpenAI started design in mid-2024 and handed the chip to the fab in November 2025. That is nine months from first design to manufacturing handoff, and 16 months to tape-out — a schedule that normally takes a silicon team two to three years.

    OpenAI says its own models helped design the chip. That claim is unverifiable from the outside, but the calendar is not.

    “Jalapeño can serve more AI work per unit of power, while also returning responses more quickly,” said Richard Ho, OpenAI’s head of hardware, in comments reported by TechCrunch.

    How much faster is Jalapeño than Nvidia Blackwell?

    Across three open-weight models, Jalapeño roughly doubled Nvidia’s tokens per second per kilowatt while cutting latency by 43% to 72%. The gap widens as models get larger. On DeepSeek R1 670B, Jalapeño returned a first response in 1.65 seconds against GB300’s 5.99 seconds.

    Here are the SemiAnalysis InferenceX results as reported by ForkLog:

    Model Jalapeño (mixed TPS/kW) Nvidia system Nvidia (mixed TPS/kW) Jalapeño latency Nvidia latency
    GPT-OSS 120B 85,448 GB200 44,960 1.03s 1.80s
    DeepSeek R1 670B 19,641 GB300 11,781 1.65s 5.99s
    Kimi K2.5 1T 18,195 GB300 11,862 1.56s 5.31s

    On single-user throughput, Jalapeño hit roughly 1,400 tokens per second on GPT-OSS 120B and over 700 tokens per second on DeepSeek R1 670B.

    The aggregate claims are wider still: 1.7x to 3.6x lower end-to-end latency and 2.1x to 4.1x higher performance on interactive workloads, per The Decoder. At matched decoding speeds, The Decoder reported token-throughput-per-kilowatt advantages of 54x to 104x — a number that only makes sense in the narrow regime where GPU batching collapses.

    What SemiAnalysis actually said

    “Usually first generation chips aren’t competitive, but OpenAI is beating Nvidia Blackwell and even Rubin,” SemiAnalysis CEO Dylan Patel said, per The Decoder.

    That is a strong endorsement from an analyst house that sells research to the same hyperscalers buying Nvidia racks. Take it seriously. Take it with salt.

    Why does performance per watt decide who wins?

    Because power, not silicon, is the binding constraint on AI buildouts in 2026. Data center operators are queuing for grid interconnects measured in years. If a chip does the same work at half the watts, the same substation serves twice the revenue.

    That math is why custom ASICs keep appearing. Every watt saved on inference is a watt available for a paying customer, and inference is now the majority of frontier-lab compute spend.

    OpenAI CFO Sarah Friar framed it in cost terms: custom chips give the company “greater control over inference costs” and let it match hardware to specific tasks. Friar also said the chip “complements” existing partnerships rather than replacing them — corporate language for we are still buying your GPUs, please keep taking our calls.

    We covered the same power-and-memory squeeze from the supply side in our piece on the Nvidia AI server price hike, and the economics of fast inference in Cerebras vs Groq.

    What does this do to Nvidia’s margins?

    Nothing this quarter. Nvidia reported Q2 fiscal 2027 revenue of $96.22 billion on August 26, beating the $92.07 billion consensus, with data center revenue of $89.02 billion — up 117% year over year, according to 24/7 Wall St. EPS came in at $2.22 against a $2.09 estimate.

    Guidance was louder than the beat. Nvidia guided Q3 to $108 billion plus or minus 2%, with non-GAAP gross margins near 74% and no China data center compute revenue assumed.

    “AI has reached its inflection point. It’s doing useful work. Its tokens are productive and profitable. Now, compute is revenue,” CEO Jensen Huang said on the call.

    Nvidia also disclosed supply commitments of $279 billion, largely for Vera Rubin memory. That is a company buying ahead, not one bracing for demand loss.

    The threat is 2028, not 2026

    Custom silicon does not eat Nvidia’s revenue. It eats Nvidia’s pricing power. A 74% gross margin exists because there is no substitute at scale. Jalapeño is the first credible substitute built by Nvidia’s single largest customer.

    NVDA closed at $213.05 before the print, down 3.04% on the week and up 14.37% year to date, per 24/7 Wall St. The stock has fallen after four of its last five earnings reports despite beating consensus three quarters running.

    Who wins and who loses financially?

    Broadcom is the clearest winner. It gets ASIC design revenue, a marquee reference customer, and validation that its custom-silicon business can beat the merchant-GPU incumbent on a first attempt. Nvidia is the clearest medium-term loser, though the damage lands in 2028 pricing, not 2026 volume.

    • Broadcom — books high-margin custom ASIC revenue and proves the model. We covered its financing appetite in the Broadcom AI debt deal.
    • TSMC — wins either way. N3P wafers are N3P wafers, whether the logo says Nvidia or OpenAI.
    • HBM suppliers — Jalapeño uses HBM4 at 15.4 TB/s. More custom chips means more high-bandwidth memory demand, not less.
    • OpenAI — gains leverage in every future GPU negotiation, which may be worth more than the chip itself. Its Nvidia relationship already shifted once, as we noted when Nvidia cut its OpenAI data center guarantee.
    • Nvidia — keeps the volume through 2027, then defends 74% margins against a credible in-house alternative.
    • Second-tier inference clouds — squeezed hardest. They rent GPUs at market rates and cannot design their own.

    What’s the catch with the Jalapeño benchmarks?

    Three catches, and they matter. Jalapeño exists as engineering samples only. Rubin is already shipping to customers. And the benchmark set was chosen by the chip’s owner, run on three open-weight models, with two of Nvidia’s standard optimizations absent from the comparison.

    The Decoder reported that Jalapeño lacks multi-token prediction and speculative decoding optimizations. Those are exactly the techniques that close latency gaps on GPUs. Adding them later helps Jalapeño; adding them to the comparison today would narrow the gap.

    The models tested were GPT-OSS 120B, DeepSeek R1 670B and Kimi K2.5 1T. Larger current-generation models — DeepSeek V4 Pro, Kimi K3 — were not tested at all. Neither, notably, was any GPT-5-class OpenAI frontier model, which is the workload the chip actually has to serve.

    And the deployment schedule is honest about itself: very small volumes at the end of 2026, meaningful volume in 2027. OpenAI says a second generation is in advanced development and a third is in design.

    A chip that wins benchmarks in August 2026 must still win against whatever Nvidia ships in 2027. That is a different race.

    Frequently asked questions

    Is the OpenAI Jalapeño chip available to buy?

    No. It is an internal accelerator for OpenAI’s own inference fleet, currently at engineering-sample stage. Small-volume deployment starts at the end of 2026, with wider rollout in 2027. There is no external sales channel announced.

    Who manufactures the Jalapeño chip?

    Broadcom co-developed it with OpenAI, and TSMC fabricates it on the N3P process node. The benchmarked silicon is B0 stepping, meaning at least one revision past first tape-out.

    Does Jalapeño beat Nvidia’s Rubin?

    On the perf-per-watt figures SemiAnalysis published, yes — 1.5x to 1.9x. But Rubin is shipping to paying customers now and Jalapeño is not, so the comparison is between a product and a prototype.

    Can Jalapeño train models?

    No. It is an inference-only design. OpenAI still needs GPUs for training, which is why CFO Sarah Friar described the chip as complementing rather than replacing existing supplier relationships.

    How much power does Jalapeño use?

    It is rated at 700W and reportedly sustains under 550W in operation. Nvidia’s GB200 draws 1,200W and GB300 draws 1,400W, so Jalapeño operates in roughly half the envelope.

    Did Nvidia’s earnings show any damage from custom chips?

    None yet. Data center revenue grew 117% year over year to $89.02 billion and Q3 guidance is $108 billion. Custom silicon is a 2028 margin question, not a 2026 revenue question.

    What benchmark was used?

    SemiAnalysis InferenceX, which measures mixed tokens per second per kilowatt alongside end-to-end latency. It is a third-party benchmark, but the model selection and test configuration came from the chip’s owner.

    The bottom line

    Jalapeño is the most serious first-generation AI accelerator anyone has produced, and the power numbers are the part that should worry Nvidia. Half the watts for double the tokens is not a rounding error; it is a structural argument for custom silicon at every lab large enough to fund a design team.

    But the trade here is not “sell Nvidia.” Nvidia just printed $96.22 billion in a quarter and guided to $108 billion. The trade is that Nvidia’s 74% gross margin now has an expiry date attached, and the market will start pricing that date long before 2028 arrives.

    The honest read: OpenAI has proven it can build a chip. It has not yet proven it can build ten million of them, on schedule, while Nvidia iterates annually. Benchmarks are cheap. Yield is not.

    Sources