Category: AI News

  • Claude Watermark Backlash: $100 Users Quit Over Invisible Marks

    Anthropic’s Claude watermark is now live worldwide. Every Claude model launched on or after August 2, 2026 embeds an invisible SynthID-Text mark in generated text — across the API, Claude Code and Claude Cowork. Some $100-a-month Max subscribers are canceling in protest. Anthropic says it has seen no statistically significant uptick, while posting an $11.5 billion quarter.

    The feature was announced on August 11. The technical detail landed on Friday, August 15, in an Anthropic blog post covered by TechCrunch. The cancellation screenshots started over the weekend.

    That sequence matters. Anthropic is meeting investors ahead of a possible fall IPO. It just reported its first quarter of positive adjusted operating income.

    And it chose this moment to stamp a machine-readable signature on everything its models write.

    What is the Claude watermark?

    The Claude watermark is an invisible statistical signature embedded in text the model generates. It nudges Claude toward one of several equally valid word choices, following a secret key. Readers see nothing. A detector holding the key can spot the pattern. Image and file outputs get separate provenance metadata instead.

    How the token steering works

    Anthropic uses the SynthID-Text approach that Google DeepMind published in 2024. The system biases “low-stakes” token choices — preferring “overcast” over “grey” in a sentence about cold weather.

    Where only one answer is correct, the watermark stays off. Gizmodo reported that in a prompt like “Paris is the capital of…” there is no room to steer, so nothing is embedded.

    Anthropic told TechCrunch: “Watermarking does not impact the quality of Claude’s output. To a reader, a watermarked response is indistinguishable from an unwatermarked one.” The company describes the effect on speed and token cost as “negligible.”

    What actually gets marked

    Per Anthropic’s own support documentation, embedded watermarks “apply to all generated text.” Files get signed provenance metadata following the C2PA open standard — the documentation names .svg, .png and .jpg.

    Code is the exception. Functional code demands specific tokens, so there is almost nothing to steer. TechCrunch reported that marking shows up mainly in optional elements such as comments.

    Why did Anthropic ship the Claude watermark now?

    Because a legal deadline landed. Article 50 of the EU AI Act applied from August 2, 2026, and requires providers of generative systems to mark synthetic text, audio, image and video in a machine-readable format. Anthropic’s own docs confirm models launched on or after that date “support marking at launch.”

    Legacy models get a grace period. Under the AI Omnibus agreement, systems already on the market before August 2026 have until December 2, 2026 to meet the machine-readable marking requirement.

    The part that annoyed users is the geography. Search Engine Land reported that “Anthropic said it will enable the watermarking system everywhere Claude is available, not just in Europe.”

    That is a choice, not an obligation. A US freelancer with no EU exposure now carries an EU compliance artifact in their drafts — and Anthropic gets to walk into IPO meetings as the lab that shipped transparency early. Read that alongside our coverage of Anthropic’s $2 trillion IPO positioning and the timing looks less accidental.

    Who is affected by the Claude watermark?

    Effectively everyone who touches Claude. Anthropic states the marking applies “everywhere you use Claude, including Claude Platform (API), Claude, Claude Code, Claude Cowork, and Claude Tag,” and that it also applies when supported models are accessed through AWS, Google Cloud or Microsoft Foundry. There is no documented consumer-versus-enterprise carve-out.

    SurfaceMarked?What it means in practice
    Claude appsYes — textChat output carries the signature
    Claude Platform (API)Yes — textYour product’s output is marked too
    Claude CodePartialMostly comments, not functional code
    Claude CoworkYes — text and filesC2PA metadata on generated images
    AWS / Google Cloud / Microsoft FoundryYesCloud resale does not strip the mark
    Models released before Aug 2, 2026Rolling outRetrofit due by Dec 2, 2026

    Are Claude users really canceling their subscriptions?

    Some are, loudly. Multiple subscribers posted cancellation screenshots on X citing the watermark, including math influencer John Ennis, who called it a “ridiculous watermark idea.” The tier repeatedly named in coverage is Claude Max at $100 per month.

    Anthropic’s counter is blunt. The company told Business Insider it had not yet seen a statistically significant uptick in cancellations since the announcement.

    Both things can be true. A few hundred visible cancellations is a trend on X and a rounding error on the income statement.

    The scale gap is the story. Fortune reported Anthropic’s Q2 2026 revenue at over $11.5 billion, against $787 million in Q2 2025 — at least 14-fold growth — with positive adjusted operating income. Q1 2026 was $4.73 billion.

    Consumer subscriptions are not what produced those numbers. The API is. And API customers building products in Europe want the compliance box ticked far more than they want unmarked prose.

    Can you remove the Claude watermark?

    Partly, and inconsistently. Anthropic says light editing “probably won’t remove the watermark completely,” while “a complete rewrite where every word is replaced will.” Detectability scales with text length, so short outputs may carry no usable signal at all.

    • Full rewrites strip it — every token has to change.
    • Light edits usually do not strip it.
    • Short passages may never accumulate enough signal to detect.
    • Code is barely marked, so engineering workflows are largely unaffected.
    • Passing Claude text through a second model is the obvious laundering route — and nothing in the design prevents it.

    There is a nastier edge case. TechCrunch noted the mark can appear even when Claude is only editing text a human wrote — which is exactly the scenario freelancers and consultants are worried about.

    Who wins and who loses financially?

    Anthropic wins on regulatory positioning ahead of an IPO. Detection and provenance vendors win a mandated market. The losers are the people whose pricing depends on clients believing a human wrote the words — and, potentially, Anthropic’s own consumer tier.

    The winners

    Anthropic first. Fortune reports the company filed confidentially and is working with Morgan Stanley, Goldman Sachs and JPMorgan Chase, with an autumn listing that would beat OpenAI and DeepSeek to market. “Already compliant with Article 50” is a clean line in a prospectus.

    Then the provenance stack. C2PA tooling, detection APIs and audit vendors all get a regulatory tailwind — Anthropic has said it plans to release a watermark detection API of its own.

    Platforms benefit too. Axios reported LinkedIn testing an AI detection button, Substack integrating detection technology and Snap limiting promotion of AI-generated video.

    The losers

    Freelance writers, agencies and consultants who bill human rates. A watermark that survives light editing turns “I used AI to tidy this up” into an evidentiary problem with a client.

    Competitors get a gift. Axios reports OpenAI’s watermarking efforts “primarily focus on images and audio rather than text.” For a writer choosing a $100 subscription this week, that asymmetry is the whole decision — and it lands in a market where the price war has already turned.

    There is also an open-weights escape hatch. Self-hosted models carry no vendor watermark at all, which quietly strengthens the case we made in our breakdown of open weights versus API.

    Is the Claude watermark actually reliable?

    No — and Anthropic says so itself. The company cautions that a detected mark “doesn’t prove Claude created the original ideas or wrote the original text,” and that “the absence of a detectable watermark doesn’t mean content wasn’t AI-generated.” Both directions fail.

    Read that carefully. A positive result does not prove authorship. A negative result does not prove human writing.

    What is left is a signal that Claude touched a document at some point. That is a compliance artifact, not evidence.

    The skeptical read: universities, publishers and HR departments will treat it as evidence anyway, because a machine-readable flag is easier to act on than a judgment call. Anthropic has built a tool it explicitly says cannot answer the question everyone will use it to answer.

    Claude watermark FAQ

    Can I turn the Claude watermark off?

    Anthropic’s published documentation lists no opt-out, and none of the major coverage from August 11–16 identified one. Assume it is on.

    Does it apply outside the European Union?

    Yes. Anthropic is enabling it worldwide, not only in the EU, even though Article 50 is a European obligation.

    Does it slow Claude down or cost more tokens?

    Anthropic describes the impact on speed and token cost as negligible. No independent benchmark of that claim has been published yet.

    Will it flag code I wrote with Claude Code?

    Barely. Functional code leaves almost no room for token steering, so marking concentrates in comments and other optional text.

    Which models are affected?

    Claude models launched on or after August 2, 2026 support marking at launch. Earlier models are being retrofitted during the AI Act transition period, which runs to December 2, 2026.

    Can a client or university prove I used Claude?

    Not conclusively. Anthropic states the mark shows Claude processed the text, not that Claude authored it. Expect that nuance to be ignored in practice.

    Do OpenAI and Google watermark text the same way?

    Not equivalently. Axios reports OpenAI’s watermarking focuses on images and audio rather than text. The underlying SynthID-Text method originated at Google DeepMind in 2024.

    The bottom line

    The Claude watermark is a compliance product shipped as a trust product, and the gap between those two things is where the backlash lives.

    Anthropic is not going to lose an $11.5 billion quarter over cancelled $100 subscriptions. The API business it actually runs on will benefit from being first through the Article 50 door.

    But it has handed competitors a specific, nameable reason to switch — and handed self-hosted open weights a genuine selling point that has nothing to do with price or benchmarks.

    Verdict: right call for the IPO, expensive call for the writers. If your work product depends on nobody being able to run a detector over it, this was the week your vendor choice changed.

    Sources

  • SMIC AI Chip Demand Drives First $3 Billion Quarter, Profit Up 262%

    SMIC AI chip demand pushed China’s largest foundry past $3 billion in quarterly revenue for the first time. Q2 2026 revenue hit $3.006 billion, up 36.1% year over year, while net profit attributable to owners jumped 261.7% to $479.2 million. Gross margin widened to 25.3% from 20.4%. Wafer prices rose 5.7% sequentially — and management says more increases are coming in Q3.

    How much did SMIC make in Q2 2026?

    Semiconductor Manufacturing International Corporation reported $3.006 billion in revenue for the quarter ended June 30, 2026 — a 20.0% sequential increase and 36.1% growth year over year. Net profit attributable to owners reached $479.2 million, up 261.7%. Both figures beat analyst estimates compiled by LSEG, according to Reuters.

    It is the first time SMIC has cleared $3 billion in a single quarter. Revenue was $2.51 billion in Q1 2026 and $2.21 billion in Q2 2025, per Global Times.

    The profit line moved far faster than the top line. Operating profit rose 254.5% year over year to $534.2 million. That gap is the whole story of this quarter.

    The margin story behind the revenue

    Gross profit came in at $760.6 million, a 69.1% year-over-year increase. Gross margin expanded to 25.3%, up from 20.1% in Q1 2026 and 20.4% a year earlier.

    Two things drove it: volume and price. Wafer shipments reached 2.869 million 8-inch-equivalent units, up 14.4% sequentially. Average selling price climbed 5.7% over the same period.

    Metric Q2 2025 Q1 2026 Q2 2026
    Revenue $2.21B $2.51B $3.006B
    Gross margin 20.4% 20.1% 25.3%
    Net profit (attributable) ~$132M ~$197M $479.2M
    Wafer shipments (8-inch equiv.) 2.39M 2.51M 2.869M
    Fab utilization 92.5% 93.1% 93.7%
    China share of revenue 84.1% 88.9% 90.2%

    Q2 2025 and Q1 2026 profit and shipment figures are derived from the disclosed growth rates (+261.7% and +142.7% for profit; +20.1% and +14.4% for shipments).

    Why is SMIC raising wafer prices?

    Because it can. Fab utilization hit 93.7% and monthly capacity grew just 1.7% sequentially to 1.097 million wafers. When demand outruns capacity that tightly, price becomes the only lever left. SMIC raised prices after Q1 customer negotiations and has told investors more increases land in Q3.

    The company added only 8,000 wafers per month of 12-inch capacity during the quarter, Reuters reported. That is a rounding error against a 1.1 million-wafer base.

    What Zhao Haijun actually said

    Co-CEO Zhao Haijun framed the increases as a correction rather than opportunism. “Since there’s still a big gap between industry-leading wafer prices and SMIC’s current prices, we need to negotiate with customers for fairer pricing,” he said, per Taipei Times.

    He also described where the volume came from: “The rise in shipments was driven mainly by surging AI-fueled demand for chips other than CPUs and GPUs, mostly from China-based customers.”

    That second quote deserves more attention than it got. Read it again.

    What kind of AI chips is SMIC actually making?

    Not the accelerators. SMIC’s AI exposure is in the supporting silicon around the compute — power management ICs, controllers, connectivity parts, and interface chips. US export controls still keep the company off the leading-edge nodes where AI training chips are fabricated.

    The application mix confirms it. Consumer electronics accounted for 44.2% of Q2 revenue, smartphones 16.9%, industrial and automotive 16.5%, computer and tablet 15.6%, and connectivity and IoT 6.8%.

    Per AnySilicon’s breakdown of the results, AI-supporting products, computers and tablets, and industrial and automotive applications each grew roughly 40% sequentially. Twelve-inch wafers now represent 78.2% of revenue, up from 76.1% a year ago.

    • Power management ICs — every rack of AI servers needs hundreds
    • Controllers and interface chips — the connective tissue of a datacenter
    • Connectivity and IoT silicon — 6.8% of revenue and growing
    • Industrial and automotive — 16.5%, boosted by China’s intelligent-driving push

    How does Hua Hong’s quarter compare?

    Hua Hong Grace Semiconductor posted record quarterly revenue of $717.5 million, ahead of the $702.7 million consensus. Gross margin of 16.5% beat the 14.9% forecast, though net profit of $38.6 million came in slightly below the $39.4 million estimate. Guidance calls for $770–780 million in Q3 at 16–18% margins.

    Both foundries are lifting prices. Both are running near capacity. The pattern is sector-wide, not company-specific.

    Analyst Ma Jihua told Global Times that China’s domestic semiconductor sector “is benefiting from several factors at once, including strong downstream demand, policy support and market space created by US restrictions.”

    Why this matters for the AI market and investors

    The AI trade has been priced almost entirely off the accelerator layer — Nvidia, Broadcom, the hyperscaler capex line. SMIC’s quarter is evidence that the mature-node tier underneath is also capacity-constrained and gaining pricing power. That is a cost input for everyone building AI hardware.

    It also complicates the export-control thesis. Restrictions were designed to slow China’s advanced-node capability. They did. But they also handed SMIC a protected domestic market: China went from 84.1% of revenue a year ago to 90.2% today.

    Goldman Sachs maintained a buy rating with a HK$135 price target after the print, per Global Times. SMIC shares rose 4.81% in Hong Kong.

    The skeptical read

    Three things temper the enthusiasm.

    First, the deceleration. Revenue grew 20.0% sequentially in Q2. Management guides Q3 to just 2–4%. Capacity, not demand, is the ceiling — and adding 8,000 wafers a month will not move it.

    Second, the margin gain is priced, not earned through technology. ASP rose 5.7% because customers had nowhere else to go. That is a real advantage, but it is not the same as node leadership, and it invites customers to qualify second sources.

    Third, the depreciation wave is coming. H1 capex was $3.4 billion against $3.3 billion a year prior, and full-year amortization is running around $5 billion, up roughly 30% year over year. Those costs land on future margins regardless of what prices do.

    A 90% domestic revenue concentration is a strength in a fragmenting market and a liability in a consolidating one. It is not obvious which one 2027 delivers.

    This post is analysis and reporting, not financial advice.

    What are the risks to SMIC’s momentum?

    The near-term risk is not demand — it is the arithmetic of capacity plus depreciation. SMIC can raise prices only while customers lack alternatives. Chinese fabs are expanding aggressively, and every new line that qualifies erodes the scarcity premium that produced this quarter’s 25.3% gross margin.

    1. Capacity catching up. Domestic competitors adding mature-node lines through 2027
    2. Amortization drag. ~$5 billion this year, up ~30%, hitting reported margins
    3. Customer concentration. 90.2% of revenue from a single geography
    4. Node ceiling. Export controls still bar the leading edge
    5. Price fatigue. Two consecutive increases invite qualification of second sources

    Frequently asked questions

    How much revenue did SMIC report in Q2 2026?

    $3.006 billion, up 36.1% year over year and 20.0% sequentially. It was the first quarter in company history above $3 billion.

    How much did SMIC’s profit grow?

    Net profit attributable to owners rose 261.7% year over year to $479.2 million, and 142.7% from Q1 2026.

    Is SMIC making AI accelerator chips?

    No. Zhao Haijun attributed the shipment growth to AI-driven demand for chips other than CPUs and GPUs. SMIC supplies power management, controller, and connectivity silicon that surrounds AI compute.

    Why did SMIC raise wafer prices?

    Utilization hit 93.7% with monthly capacity up only 1.7%. Zhao said there remains “a big gap between industry-leading wafer prices and SMIC’s current prices.”

    What is SMIC’s Q3 2026 guidance?

    Revenue growth of 2–4% sequentially, with gross margin between 26% and 28% — a sharp deceleration from Q2’s 20.0% sequential growth.

    How did the market react?

    SMIC shares rose 4.81% in Hong Kong. Goldman Sachs kept a buy rating with a HK$135 target, according to Global Times.

    Did Hua Hong Semiconductor report similar strength?

    Yes. Hua Hong posted record revenue of $717.5 million against a $702.7 million estimate, with gross margin of 16.5% versus a 14.9% forecast.

    The bottom line

    SMIC just proved that the AI buildout pays out well below the accelerator layer. A foundry barred from the leading edge posted a 261.7% profit increase by selling ordinary chips into an extraordinary shortage.

    The question for the next two quarters is whether that shortage is structural or a timing artifact. Guidance of 2–4% sequential growth suggests SMIC itself is not certain.

    Watch three numbers in the Q3 print: gross margin against the 26–28% guide, the China revenue share, and monthly capacity additions. If margin lands at the top of the range while capacity stays flat, pricing power is real. If capacity jumps and margin slips, this quarter was the peak.

    Sources

    Related on Wealth Engine: Qwen3.8-Max open weights and the real cost of self-hosting · DeepSeek’s price increase and the end of the AI price war · Five companies committed $650 billion to AI in a single year · Anthropic’s $2 trillion valuation and the Decart acquisition

  • Sora 2 Alternatives: Cheapest AI Video API Before the Sept 24 Sunset

    Sora 2 alternatives are now a deadline, not a preference. OpenAI kills the Sora API on September 24, 2026 — 38 days away. The cheapest replacement is Veo 3.1 Lite at $0.03 per second, video-only. For audio-native output, Veo 3.1 Lite runs $0.05. For visual quality per dollar, Kling 3.0 at $0.084 already beat Sora 2’s $0.10. Migrate now.

    OpenAI is walking away from video. That is the story buried inside a support-page sentence, and it forces every product team still calling the Sora endpoint to pick a replacement before the end of September.

    The good news for your budget: the market moved past Sora while OpenAI was deciding to leave it. The replacements are cheaper, longer, and in several cases score higher with human raters.

    Why is the Sora 2 API shutting down, and when?

    OpenAI confirmed the timeline in its own help center: the Sora web and app experiences were discontinued on April 26, 2026, and the Sora API will be discontinued on September 24, 2026. After a stated export window, OpenAI will permanently delete data associated with your Sora usage.

    What OpenAI is actually deleting

    This is not a version bump. There is no Sora 3 waiting behind it. OpenAI’s discontinuation notice offers no replacement product and instead points users to sora.chatgpt.com/sunset to export their generations.

    Unused Sora credits can be redirected toward other OpenAI services such as Codex. That is a tell. OpenAI is reallocating spend toward agents and coding, the same direction we tracked when Meta shipped Muse Spark at 4x cheaper coding economics.

    Why Sora lost

    Sora stopped appearing on public video leaderboards well before the sunset was announced. On the Artificial Analysis text-to-video arena — blind human voting — the August 2026 top three are Gemini Omni Flash at roughly 1,238–1,245 Elo, MiniMax H3 at 1,235–1,242, and ByteDance’s Seedance 2.0 at 1,220–1,225.

    Sora 2 is not in that list. It is not in the top ten. A model that is neither cheapest nor best is a model with no reason to exist.

    Which Sora 2 alternatives are cheapest per second?

    Veo 3.1 Lite is the cheapest credible option at $0.03 per second video-only and $0.05 with native audio. Kling 3.0 Standard sits at $0.084, and ByteDance’s brand-new Seedance 2.5 lists from $0.1028. Sora 2’s old $0.10 rate is now mid-pack, not competitive.

    Model Price/sec (video only) Price/sec (with audio) Notes
    Veo 3.1 Lite $0.03 $0.05 720p; cheapest audio-native option
    Runway Gen-4 Turbo $0.05 5 credits/sec, no native audio
    Kling 3.0 Standard $0.084 $0.126 720p; $0.112–$0.168 at 1080p
    Veo 3.1 Fast $0.10 $0.15 720p production tier
    Sora 2 (dying) $0.10 n/a 720p; $0.05 batch. Off Sept 24
    Seedance 2.5 from $0.1028 optional Released Aug 7, 2026; 30s clips
    Kling 3.0 Turbo $0.112 $0.56 per 5-second clip
    Runway Gen-4.5 $0.12 12 credits/sec
    FLUX 3 Video $0.17 included 20s clips, native dialogue
    Veo 3.1 Quality $0.20 $0.40 720p/1080p flagship
    Sora 2 Pro (dying) $0.30–$0.70 n/a 720p to 1080p
    Sources: CometAPI pricing index (updated Aug 16, 2026), CostGoat Veo and Sora calculators, OpenRouter model pages, Renderful Kling pricing.

    Read that table as a verdict, not a menu. Sora 2 was charging $0.10 per second for a model that did not rank, while Google was selling audio-native video at half the price.

    Is the cheapest AI video API also the best?

    No — but the gap is smaller than the price gap. Veo 3.1 ranks around #11 overall on the Artificial Analysis arena despite being the cheapest audio-native option. Kling 3.0 1080p Pro sits at 1,107 Elo and Kling 3.0 720p at 1,099, both inside the top ten while costing under $0.13 per second.

    What the arena actually measures

    The Artificial Analysis leaderboard is blind human preference voting, not a technical benchmark. It rewards prompt adherence and perceived realism. It does not measure API reliability, rate limits, or how a model handles your specific reference images.

    Rankings also move weekly. Alibaba’s Wan2.7 sits at 1,158 Elo and Skywork’s SkyReels V4 at 1,103 — close enough that a single release reshuffles the middle of the board.

    Which Sora 2 alternative should you pick for your use case?

    Match the model to the job, not the leaderboard. Audio-synced marketing video goes to Veo 3.1. Multi-shot narrative sequences go to Kling 3.0. Image-anchored generation goes to Seedance. High-volume social clips go to Veo 3.1 Lite, where the per-second price is the whole argument.

    Your Sora 2 use case Migrate to Cost for a 10s clip Why
    High-volume social clips Veo 3.1 Lite $0.50 with audio Half of Sora 2’s rate, audio included
    Ads and audio cinematics Veo 3.1 Quality $4.00 with audio Always-on audio, 4K available
    Multi-shot storytelling Kling 3.0 $0.84–$1.26 Native support for up to 6 labeled shots
    Image-to-video fidelity Seedance 2.5 ~$1.03 Up to 50 reference assets; 30s single takes
    Long-form dialogue scenes FLUX 3 Video $1.70 20s clips with native dialogue
    Cheapest possible pipeline Veo 3.1 Lite (no audio) $0.30 $0.03/sec is the floor right now
    Batch jobs you ran overnight Veo 3.1 Lite $0.30–$0.50 Matches Sora 2’s $0.05 batch rate
    Clip costs calculated from the per-second rates in the table above.

    How much does migrating off Sora 2 actually cost?

    For most teams, migration is a price cut. At 1,000 ten-second clips per month — a modest content pipeline — Sora 2 at $0.10 per second cost $1,000. Veo 3.1 Lite with audio costs $500 for the same volume. The exceptions are batch users and Sora 2 Pro users.

    • Sora 2 standard, 1,000 clips × 10s: $1,000/month. Veo 3.1 Lite: $500. You save $6,000 a year.
    • Sora 2 batch at $0.05/sec: $500/month. Veo 3.1 Lite with audio matches it exactly at $0.05 — a wash, and you gain audio.
    • Sora 2 Pro 1080p at $0.70/sec: $7,000/month. Veo 3.1 Quality with audio at $0.40: $4,000. That is a 43% cut.
    • Kling 3.0 Standard route: $840/month, plus a separate audio step that Veo bundles for free.
    • Seedance 2.5 route: roughly $1,028/month — slightly above Sora 2, but you get 30-second single takes instead of short clips.

    The engineering cost is the real line item. Budget two to four days of developer time for endpoint changes, prompt re-tuning, and regression checks on your existing library.

    Is Veo 3.1 Quality worth 4x the price of Kling 3.0?

    Only if audio is non-negotiable. Veo 3.1 Quality with audio costs $0.40 per second against Kling 3.0 Standard’s $0.084 — a 4.8x premium. Veo is the only major model shipping native audio inside the video output. Every alternative needs a separate audio generation step.

    That separate step is not free. It adds latency, a second vendor, and a sync problem. If you are producing narrated ads, Veo’s bundled audio is worth the premium.

    If you are producing silent B-roll, product loops, or background video, paying $0.40 for audio you will mute is the single worst decision available in this market. Use Veo 3.1 Lite at $0.03 and keep the difference.

    How do you migrate off Sora 2 without breaking production?

    Export first, then swap endpoints, then re-tune prompts. The export window is the only irreversible deadline — OpenAI deletes Sora-associated data after it closes. Everything else can be fixed after September 24. Losing your generation history cannot.

    1. Export today. Pull your full Sora library from the sunset page before the window closes. This takes an hour and cannot be undone later.
    2. Inventory your prompts. Sora prompts do not transfer cleanly. Veo and Kling weight camera language and shot structure differently.
    3. Run a 20-clip bake-off. Generate the same 20 prompts on Veo 3.1 Lite, Kling 3.0, and Seedance 2.5. At these prices the whole test costs under $30.
    4. Check audio separately. If you pick Kling or Seedance, price your audio vendor into the per-second math before you commit.
    5. Keep a second provider wired. Sora’s shutdown is the argument for never having one video vendor again.

    Frequently asked questions about Sora 2 alternatives

    When exactly does the Sora 2 API stop working?

    September 24, 2026, per OpenAI’s own help center. The consumer app already shut down on April 26, 2026.

    Is there a Sora 3 coming?

    OpenAI has announced no replacement video model. Its discontinuation notice offers no successor product and suggests redirecting unused credits to services like Codex.

    What is the cheapest Sora 2 alternative?

    Veo 3.1 Lite at $0.03 per second video-only, or $0.05 with native audio. That is half of Sora 2’s $0.10 standard rate.

    Which AI video model ranks highest right now?

    Gemini Omni Flash leads the Artificial Analysis text-to-video arena at roughly 1,238–1,245 Elo, followed by MiniMax H3 and ByteDance’s Seedance 2.0 at about 1,220–1,225.

    Is Seedance 2.5 worth switching to?

    If you work from reference images. Released August 7, 2026, it accepts up to 50 image, video, and audio reference assets and generates 30-second single takes, listed from $0.1028 per second on OpenRouter.

    Do I lose my old Sora videos?

    Yes, unless you export them. OpenAI states it will permanently delete data associated with your Sora usage after the export window closes.

    Which model has native audio built in?

    Veo 3.1 across all tiers, Kling 3.0 at a $0.042 per-second premium, and FLUX 3 Video, which ships native dialogue in its 20-second clips.

    The bottom line

    Migrate to Veo 3.1 Lite. At $0.05 per second with native audio, it is half of what Sora 2 charged without audio, and it matches Sora 2’s batch rate at full standard pricing. For 90% of teams that is the answer, and the migration pays for itself in the first month.

    Choose Kling 3.0 instead if you need multi-shot narrative control and already own an audio pipeline. Choose Seedance 2.5 if your workflow is image-anchored and 30-second takes matter more than $0.02 per second.

    Choose Veo 3.1 Quality only when a client is paying for broadcast-grade audio cinematics. At $0.40 per second it is 13x the price of the Lite tier.

    The larger lesson is the one we flagged when DeepSeek raised prices up to 1,100% overnight: inference pricing is not a stable input. Sora went from flagship to deleted in sixteen months. Build your stack so the next sunset costs you a config change, not a quarter.

    Related reading: Gemini 3.7 Flash vs Claude Sonnet 5 on cost per coding point and the real cost of self-hosting open-weights models.

    Sources


  • Muse Spark vs Claude Opus 5: Meta’s 4x Cheaper Coding Model

    Muse Spark vs Claude Opus 5 comes down to one number: price. Meta’s coding model lists at $1.25 per million input tokens against Anthropic’s $5, and Muse Spark 1.1 currently sits at the top of Scale’s SWE-bench Pro public leaderboard with 61.5%. Claude Opus 5 is still the stronger generalist. For high-volume agentic coding, Meta is now the cheaper buy — by a wide margin.

    What is Muse Spark, and why does it matter right now?

    Muse Spark is Meta’s first closed, paid model line, aimed squarely at coding. Muse Spark 1.1 launched commercially on July 9, 2026 at $1.25 per million input tokens and $4.25 per million output tokens, with $20 in free credits for new accounts, according to MarketScale.

    Muse Spark 1.2 followed on August 5, 2026, with a 1M-token context window and the same headline rates.

    A day later, Meta shipped Muse Code, an agent built on 1.2 that Forbes described as taking on “whole engineering jobs across large repositories, planning the change, writing the code, and checking the result.”

    That is the competitive set: not a chatbot, an autonomous coding worker priced to undercut everyone.

    The contributor tier is the actual product

    Meta runs two price lists. The standard tier is $1.25 in / $4.25 out. The contributor tier is $0.10 in / $0.20 out — roughly 12x to 21x cheaper, per Forbes — and the price of entry is letting Meta train on your prompts and completions.

    Cached input on the contributor tier drops to $0.01 per million tokens, according to MetaTalks.

    This is Meta’s old playbook repriced for developers. The product is the data. The discount is what they will pay for it.

    Which is better for coding, Muse Spark or Claude Opus 5?

    On the hardest public agentic benchmark, Muse Spark leads. On broad capability, Claude wins. Scale’s SWE-bench Pro public leaderboard puts Muse Spark 1.1 first at 61.5%, ahead of gpt-5.4 (xHigh) at 59.1% and claude-opus-4-6 (thinking) at 51.9%. Meta’s own numbers tell a less flattering story elsewhere.

    What SWE-bench Pro actually measures

    SWE-bench Pro is Scale’s contamination-resistant successor to SWE-bench Verified. It spans 1,865 tasks across 41 repositories — 731 public instances from GPL-licensed code, 276 from private startup codebases, and 858 held out entirely.

    The metric is resolve rate: the patch must fix the issue and not break existing tests.

    The lead holds on the harder split. On the private dataset, Muse Spark 1.1 scores 51.5%, claude-opus-4-6 (thinking) 47.1%, and gpt-5.4 (xHigh) 43.4%.

    Where Claude still wins

    Meta’s own comparison is the tell. On Meta’s internal coding benchmark, Muse Spark 1.2 scores 70.6% against Claude Opus 5’s 79.4% — a gap Meta published itself.

    On Terminal-Bench 2.1, Artificial Analysis ranks Claude Opus 5 (Adaptive Reasoning, Max Effort) at 89.1%, behind GPT-5.6 Sol (xhigh) at 89.5% and ahead of Grok 4.6 (high) at 88.4%. Muse Spark 1.2’s published Terminal-Bench 2.1 figure is 82.9%.

    Read that honestly: Meta is roughly 6 to 9 points behind the frontier on capability, and roughly 4x cheaper on input tokens. That is the entire trade.

    Muse Spark vs Claude Opus 5: how much does each cost?

    Claude Opus 5 lists at $5 per million input tokens and $25 output on Anthropic’s official pricing page. Muse Spark 1.2 lists at $1.25 and $4.25. That is 4x on input and 5.9x on output. On the contributor tier the gap widens to 50x and 125x.

    Price and spec comparison

    Model Input / 1M Output / 1M Context SWE-bench Pro (public)
    Muse Spark 1.2 (standard) $1.25 $4.25 1M 61.5% (v1.1)
    Muse Spark (contributor) $0.10 $0.20 1M 61.5% (v1.1)
    Claude Opus 5 $5.00 $25.00 1M (Opus-class) Not yet listed
    Claude Opus 4.6 $5.00 $25.00 1M (beta) 51.9%
    GPT-5.6 Sol $5.00 $30.00 1.05M Not yet listed
    GPT-5.6 Terra $2.00 $12.00 Not yet listed
    GPT-5.4 $2.50 $15.00 1M+ 59.1% (xHigh)
    Gemini 3.1 Pro Preview $2.00 (under 200k) $12.00 (under 200k) 1M 46.1%

    Sources: vendor pricing pages and Scale’s public leaderboard, retrieved August 16, 2026.

    Cost per resolved task

    Benchmarks without a price tag are marketing. Here is the arithmetic that matters.

    Assume one long-horizon agentic task burns 500,000 input tokens and 50,000 output tokens — realistic for the uncapped, 250-turn runs Scale used for its top entries. Divide the token cost by the model’s resolve rate to get the expected cost of one successfully resolved issue.

    • Muse Spark, contributor tier — $0.06 per attempt ÷ 61.5% = $0.10 per resolved task
    • Muse Spark 1.2, standard — $0.84 ÷ 61.5% = $1.37
    • GPT-5.4 (xHigh) — $2.00 ÷ 59.1% = $3.38
    • Gemini 3.1 Pro — $2.90 ÷ 46.1% = $6.29 (long-context rate: $4 / $18 above 200k tokens)
    • Claude Opus 4.6 — $3.75 ÷ 51.9% = $7.23

    Our calculation, using the listed rates above. At standard pricing Meta is 5.3x cheaper per resolved task than Opus 4.6. On the contributor tier it is 72x cheaper.

    Run 10,000 agentic tasks a month and that is roughly $13,700 on Muse Spark standard versus $72,300 on Opus. The same workload on the contributor tier costs about $1,000 — and Meta keeps your codebase patterns.

    Is Muse Spark worth it in 2026?

    For high-volume, well-specified, repetitive engineering work: yes, decisively. For novel architecture, security-sensitive code, or anything where a wrong patch is expensive, no. The 6-to-9-point capability gap against Claude Opus 5 is small in a benchmark table and large in a production incident.

    Which model should you pick?

    Use case Pick Why
    Bulk refactors, test generation, dependency bumps Muse Spark 1.2 (standard) 61.5% resolve rate at $1.37 per resolved task
    Open-source or non-proprietary code Muse Spark (contributor) $0.10 per resolved task; data sharing costs you nothing
    Proprietary IP, regulated or security-critical code Claude Opus 5 79.4% on Meta’s own benchmark; no training-data trade
    Terminal-heavy and sysadmin automation GPT-5.6 Sol or Claude Opus 5 89.5% and 89.1% on Terminal-Bench 2.1
    Cost-capped agent fleets at scale GPT-5.6 Luna or Terra Luna lists at $0.20 / $1.20 per million tokens
    Long-context repo analysis on a budget Gemini 3.1 Pro (batch) Batch mode halves rates to $1.00 / $6.00

    One caution on the benchmark itself. Forbes noted Meta published its results “as images without methodology documentation,” which is a reason to weight Scale’s independent leaderboard over Meta’s slides.

    What about GPT-5.6 and Gemini 3.1 Pro?

    OpenAI answered the price war directly. Its July 30, 2026 pricing post introduced Luna at $0.20 / $1.20 and Terra at $2.00 / $12.00 per million tokens, claiming Luna beats Fable 5 on Agents’ Last Exam at “an estimated cost per task nearly 99% lower.”

    Google sits in the middle. Gemini 3.1 Pro Preview is $2.00 / $12.00 under 200k tokens, rising to $4.00 / $18.00 above it — and 46.1% on SWE-bench Pro public is the weakest score of the frontier group.

    The pattern across all three: capability is converging, price is diverging. We saw the same dynamic when DeepSeek raised prices up to 1,100% overnight and when Qwen3.8-Max open weights forced a self-hosting math check.

    Frequently asked questions

    Is Muse Spark open weights?

    No. Muse Spark is Meta’s first closed, commercial model line — a break from the open Llama releases. Access is API-only, and Muse Spark 1.1 launched behind a waitlist.

    What is the catch with the $0.10 contributor tier?

    Meta may use your prompts and responses to improve its products. If your prompts contain proprietary source code, customer data, or unreleased product logic, the discount is not a discount.

    Does Muse Spark beat Claude Opus 5?

    Not on capability. Meta’s own published comparison shows Muse Spark 1.2 at 70.6% versus Claude Opus 5 at 79.4%. Muse Spark 1.1 does lead Scale’s SWE-bench Pro public leaderboard at 61.5%, but Opus 5 is not yet listed there.

    How big is Muse Spark’s context window?

    Muse Spark 1.2 ships with 1M tokens. Claude Opus 4.6 also offers 1M in beta, and GPT-5.6 Sol lists 1,050,000 tokens with 128k maximum output.

    Why does SWE-bench Pro matter more than SWE-bench Verified?

    SWE-bench Pro was built to resist contamination, using GPL-copyleft and private proprietary repositories. Claude Opus 4.6 scores 80.3% on SWE-bench Verified but only 51.9% on SWE-bench Pro public — the gap is the point.

    What is Muse Code?

    Muse Code is Meta’s agent product, in beta, running on Muse Spark 1.2. It handles multi-step engineering jobs across large repositories rather than single-file completions.

    Which model is cheapest per resolved coding task?

    Muse Spark on the contributor tier, at roughly $0.10 by our calculation. On standard pricing it is about $1.37, still the cheapest of the frontier group.

    The bottom line

    Buy Muse Spark 1.2 for volume, keep Claude Opus 5 for judgment. If your agent workload is high-throughput and your code is not the crown jewels, Meta’s standard tier delivers roughly 5x more resolved tasks per dollar than Opus — and the contributor tier turns that into 72x if you are willing to be training data.

    If a wrong patch costs more than a few hundred dollars to clean up, the 9-point capability gap to Opus 5 erases the savings on the first bad merge.

    The financial read is simpler still. Meta is not selling inference; it is buying developer telemetry at a 90% discount, and OpenAI’s Luna tier says it will not let that go uncontested. Model prices are heading toward zero. The margin is moving to whoever owns the agent harness — which is exactly why Cognition’s valuation hit $40 billion, and why cost per coding point is now the only benchmark that pays.

    Sources

  • Cognition $40 Billion Valuation: Up 54% in Just 11 Weeks

    Cognition is in talks to raise more than $1 billion at a valuation of at least $40 billion, Bloomberg reported on August 12, 2026. The Cognition $40 billion valuation would sit 54% above the $26 billion post-money price it closed on May 27 — just 11 weeks earlier. Annualized revenue has roughly doubled to near $1 billion in that window.

    The AI coding market has produced some fast repricing cycles. This one is close to a record.

    Cognition, the New York company behind the Devin coding agent, announced a $1 billion round at a $25 billion pre-money valuation on May 27, 2026, according to TechCrunch. Seventy-seven days later, Bloomberg reported the company back in the market at $40 billion or more.

    What is the Cognition $40 billion valuation round?

    Cognition is negotiating a new financing of more than $1 billion at a valuation of at least $40 billion, Bloomberg reported on August 12, 2026. Terms are not final and no lead investor has been named publicly. The trigger investors are pointing to is revenue: an annualized run rate approaching $1 billion.

    What we know about the terms

    • Round size: more than $1 billion, per Bloomberg.
    • Valuation: at least $40 billion — the floor, not a confirmed clearing price.
    • Status: talks. No signed term sheet has been reported.
    • Lead investor: not disclosed.
    • Existing backers: Founders Fund, General Catalyst, Lux Capital, 8VC and Khosla Ventures, per Dealroom.

    That last point matters. Every reported figure here traces back to a single Bloomberg story sourced to people familiar with the discussions. Cognition has not published anything.

    How fast did Cognition’s valuation actually climb?

    Cognition went from a $10.2 billion post-money valuation in September 2025 to a reported $40 billion floor in August 2026 — roughly 4x in 11 months. The steepest leg was the most recent: $26 billion to $40 billion in 11 weeks, without a product launch or acquisition in between.

    Date Event Amount raised Valuation
    Apr 2024 Series B, led by Founders Fund $175M $2B
    Mar 2025 Series C, led by 8VC Not disclosed $4B
    Jul 2025 Acquires Windsurf (agentic IDE)
    Sep 8, 2025 Round led by Founders Fund $400M+ $10.2B post
    May 27, 2026 Led by Lux, General Catalyst, 8VC $1B $26B post
    Aug 12, 2026 Reported talks (unsigned) $1B+ $40B+ floor
    Sources: Cognition company blog (Sep 2025), TechCrunch (May 2026), Bloomberg (Aug 2026). Pre-2025 rows per Cognition’s publicly documented funding history.

    The September 2025 round was led by Founders Fund at a $10.2 billion post-money valuation on more than $400 million raised, Cognition disclosed at the time.

    In that same post the company said Devin grew from $1 million in ARR in September 2024 to $73 million by June 2025, and that total net burn across the company’s history had stayed under $20 million.

    Does the revenue justify a $40 billion valuation?

    On the multiple, the new price is cheaper than the last one. At the May round, $26 billion against $492 million of annualized run-rate revenue was roughly 53x. At $40 billion against a run rate nearing $1 billion, it is roughly 40x. The price went up. The multiple came down.

    The multiple math

    TechCrunch reported Cognition was at $492 million ARR when the May round closed, with enterprise usage of Devin growing 50% month-over-month for six consecutive months.

    Double that base and you land near $1 billion — which is exactly the figure Bloomberg’s sources cite. The story is internally consistent, which is not the same as verified.

    Here is the skeptical read. “Annualized run rate” is one month multiplied by twelve. It is not booked revenue, it is not contracted, and at 50% month-over-month growth the number is dominated by whatever the single most recent month did.

    A company compounding that fast has an ARR figure that flatters it on the way up and punishes it the moment growth flattens. Nobody outside the round has seen net revenue retention, gross margin, or churn.

    Who else is competing for AI coding dollars?

    Cognition is not the largest asset in the category, but on revenue multiple it is the more expensive one. Cursor maker Anysphere was in talks at a $50 billion pre-money valuation on roughly $2 billion of annualized revenue as of February 2026 — about 25x, according to TechCrunch.

    That comparison cuts against the enthusiasm. At 40x, Cognition is asking investors to pay roughly $15 more of valuation for every dollar of run-rate revenue than its larger rival commanded four months ago.

    Cognition’s differentiator is positioning. Founder and CEO Scott Wu has framed Devin as a tool for “long-tail grunt-work” — legacy migrations, dependency updates — rather than a headcount replacement, per TechCrunch.

    Reported customers include Goldman Sachs, Citi, Mercedes-Benz, NASA and Santander.

    Why this matters for the AI market and investors

    Two things are happening at once, and they point in opposite directions.

    The first is that AI coding is now the clearest revenue engine in applied AI. Cognition went from $73 million ARR in June 2025 to a reported ~$1 billion 14 months later. Cursor is forecasting more than $6 billion by the end of 2026, per TechCrunch. These are not pilot budgets.

    The second is that private marks are moving faster than the businesses under them. An 11-week, 54% step-up on an unsigned round is a liquidity signal as much as a fundamentals signal — capital is chasing a small number of category leaders, and price is how it competes for allocation.

    Both can be true. The category is real and the marks are being set in a seller’s market. Meanwhile the underlying economics are still being repriced downward elsewhere in the stack, as the end of the AI price war showed this week.

    For anyone tracking exposure through secondaries or crossover funds: the marks here are set by a handful of participants in an unsigned negotiation. This post is reporting and analysis, not financial advice.

    Frequently asked questions

    Has Cognition confirmed the $40 billion valuation?

    No. As of August 15, 2026, the figure comes from Bloomberg reporting sourced to people familiar with the talks. Cognition has not issued a statement and no term sheet has been reported as signed.

    What was Cognition’s previous valuation?

    $26 billion post-money, on a $25 billion pre-money valuation, from the $1 billion round announced May 27, 2026 and led by Lux Capital, General Catalyst and 8VC, per TechCrunch.

    How much revenue does Cognition have?

    An annualized run rate approaching $1 billion, according to Bloomberg. The last independently reported figure was $492 million ARR in May 2026. Run rate is an annualized snapshot, not booked annual revenue.

    What does Cognition actually sell?

    Devin, an autonomous coding agent for engineering work, plus Windsurf, the agentic IDE Cognition acquired in July 2025. Enterprise deployment is the revenue driver.

    Who are Cognition’s investors?

    Founders Fund, General Catalyst, Lux Capital, 8VC, Khosla Ventures and Pear VC are among the disclosed backers, per Dealroom. The lead on the current round has not been reported.

    Is a 40x revenue multiple normal for AI startups?

    It is high but not an outlier in this category in 2026. Cognition’s own May round priced at roughly 53x. Anysphere’s April talks implied roughly 25x. Multiples in AI coding have been compressing as revenue scales.

    When would the round close?

    Unknown. No timeline has been reported. Cognition’s last two rounds were announced roughly eight months apart, then 11 weeks apart.

    The bottom line

    Cognition is asking the market to reprice it 54% higher on the strength of one metric moving in one direction for one quarter. The revenue growth appears real — $492 million to near $1 billion in 11 weeks is not a rounding error, and the multiple compression from 53x to 40x means the price is at least growing slower than the business.

    What to watch next: whether a named lead investor emerges, whether the final valuation clears the $40 billion floor or lands above it, and whether Cognition discloses anything beyond run rate. The company has published detailed revenue history before. If this round closes without that disclosure, that silence is the story.

    Not financial advice. Figures reported here reflect public sources as of August 15, 2026.

    Sources


  • Qwen3.8-Max Open Weights vs API: The Real Cost of Self-Hosting

    Qwen3.8-Max open weights are the biggest open release of 2026 — and the wrong choice for almost everyone. The 4.89 TB checkpoint is text-only, ships under a custom license with a $50 million revenue gate, and only beats the $2/$6 API somewhere north of six billion tokens a month. Below that, rent. Above it, talk to Alibaba’s lawyers first.

    Alibaba did something no Western lab has done this year: it put a 2.4-trillion-parameter frontier model on Hugging Face and told everyone to help themselves.

    Then it attached a license that quietly taxes anyone who succeeds with it.

    This is the deep dive on what the Qwen3.8-Max open weights actually cost to run, how they compare to the other trillion-scale open models, and the exact point where downloading beats paying.

    What exactly did Alibaba release with the Qwen3.8-Max open weights?

    Alibaba published the checkpoint to Hugging Face as Qwen/Qwen3.8-2.4T-A95B on August 8, 2026 — 224 files totaling roughly 4.89 TB, with an FP8 sibling repo, according to Digital Applied’s release checklist. It is a sparse Mixture-of-Experts model: 2.4 trillion total parameters, 95 billion active per token.

    That activation rate — about 4% — is the whole trick. You pay for a trillion-scale model’s quality while doing inference math on something closer to a 95B dense model.

    The announcement moved real money. Alibaba stock jumped 7% in Hong Kong and 4.5% on the NYSE on the news, Forkast reported.

    What’s missing compared to the hosted API

    The download is not the product Alibaba sells. The open checkpoint is text-only and thinking-mode only — no vision, and not the native 1M-token context the paid Qwen3.8-Max API advertises.

    Developers noticed within hours. The top thread on the model’s Hugging Face discussion board is titled “Huge disappointment,” pointing out that Alibaba’s launch post gave no hint the weights would be a stripped build. The promised smaller Qwen3.8-27B checkpoint still has not shipped.

    • In the download: 2.4T/95B MoE, text in, text out, thinking mode.
    • API only: vision and video input, 1M-token context (991K effective input cap), implicit prompt caching.
    • Still missing: the 27B variant, official deployment guidance, day-one community quantizations.

    How much does the Qwen3.8-Max API cost?

    List pricing is $2.00 per million input tokens and $6.00 per million output tokens, with cached input at $0.25 per million. That is the number every self-hosting calculation has to beat, and it is aggressive for a model in this weight class.

    Those figures are confirmed on OpenRouter’s Qwen3.8 Max listing, which also shows the 1M-token context window and a 131,072-token maximum output. Digital Applied notes the same rates were absent from Alibaba’s own Model Studio pricing page at launch — a reminder to check before you budget.

    For context on how fast this market moves: DeepSeek raised its own API prices by up to 1,100% overnight this month, which we covered in our breakdown of the collapsing AI price war. Cheap inference is not a permanent condition.

    What does self-hosting Qwen3.8-Max open weights actually cost?

    Far more than most teams assume. The FP8 checkpoint is roughly half the 4.89 TB full-precision drop — call it 2.4 TB of weights before you allocate a single byte to KV cache. That does not fit on one eight-GPU node once you leave room for long contexts.

    Start with the hardware rate. GetDeploying tracks 28 cloud providers offering NVIDIA B200 capacity: $3.35/hr at the cheapest reserved rate, $7.00/hr average on-demand, $3.83/hr average spot.

    The break-even math

    Take the friendliest possible case — a single eight-GPU B200 node at the cheapest reserved rate of $3.35/hr. That is $26.80/hr, or about $19,600 per month running continuously.

    Now blend the API price. At a 3:1 input-to-output ratio, Qwen3.8-Max costs $3.00 per million tokens blended. Divide $19,600 by $3.00 and you get roughly 6.5 billion tokens per month before the box is cheaper than the API.

    That is about 215 million tokens a day, every day, with zero idle time.

    ScenarioHourly rateMonthly cost (730 hrs)Break-even vs API
    1 node, reserved ($3.35/GPU-hr)$26.80~$19,600~6.5B tokens/mo
    1 node, on-demand avg ($7.00/GPU-hr)$56.00~$40,900~13.6B tokens/mo
    2 nodes, reserved (realistic minimum)$53.60~$39,100~13B tokens/mo
    2 nodes, on-demand avg$112.00~$81,800~27B tokens/mo
    GPU rates via GetDeploying; monthly figures and break-even points are our calculation at a $3.00/Mtok blended API price. Excludes engineering salaries, networking, and failed-run overhead.

    And that table is generous. It assumes 100% utilization, no redundancy node, and no one on payroll keeping the cluster alive. Add a single infrastructure engineer and the real break-even moves past 20 billion tokens a month.

    Does the Qwen3.8-Max license let you build a business on it?

    Only up to a point — and the point is $50 million. Alibaba abandoned the Apache 2.0 license used for earlier Qwen generations in favor of a bespoke qwen3.8-max license that forces large commercial users into a separate negotiation.

    Per Forkast’s analysis, any business operating as Model-as-a-Service or an “AI Work Assistant” with aggregate revenue above $50 million in any consecutive 12-month period must negotiate a commercial license. MaaS is defined broadly: any third-party access to inference or fine-tuning where the provider controls inputs or parameters.

    Read that structure carefully. It is a safe harbor for startups and a toll booth for anyone who scales.

    The strategic logic is obvious once you see it. Alibaba wants the distribution that open weights buy, without letting a competing inference layer get rich on top of its research budget. It is platform protection dressed as generosity — the same instinct behind Anthropic’s spending we analyzed in its $2 trillion valuation story, pointed a different direction.

    How do Qwen3.8-Max open weights compare to Kimi K3, DeepSeek V4 Pro and GLM-5.2?

    Qwen has the most parameters and the most restrictive license. DeepSeek V4 Pro has the best coding scores and by far the cheapest API. GLM-5.2 is the smallest and easiest to actually serve. If licensing matters to you, Qwen is the weakest of the four.

    ModelParams (total / active)LicenseAPI price (in / out per Mtok)Headline benchmark
    Qwen3.8-Max2.4T / 95BCustom, $50M revenue gate$2.00 / $6.00GPQA Diamond 92.6
    Kimi K32.8T / not disclosedModified MIT$3.00 / $15.00GPQA-Diamond 93.5
    DeepSeek V4 Pro1.6T / 49BMIT$0.435 / $0.87SWE-bench 80.6%
    GLM-5.2744B / ~40BMIT$1.40 / $4.40SWE-bench Pro 62.1
    Qwen figures via Digital Applied and OpenRouter; K3, V4 Pro and GLM-5.2 via MarkTechPost. Benchmark numbers come from different harnesses and are not directly comparable.

    Where Qwen wins, and where it loses badly

    Qwen3.8-Max leads on agentic and research tasks: 86.1 on OSWorld-Verified and 93.0 on PaperBench, ahead of GPT-5.6 Sol’s 90.5, per Digital Applied’s benchmark roundup.

    It loses on code. Qwen scores 67.7 on SWE-bench Pro against Fable 5’s 80.0, and 56.6 on DeepSWE 1.1 against GPT-5.6 Sol’s 73.0. DeepSeek V4 Pro’s 80.6% on SWE-bench beats Qwen outright while costing roughly a fifth as much on the API.

    The cost gap is the story. MarkTechPost’s Artificial Analysis blended cost-per-task figures put DeepSeek V4 Pro at $0.04, GLM-5.2 at $0.32 and Kimi K3 at $0.94. Paying 20x for a few benchmark points is a decision, not a default — the same trap we flagged in Gemini 3.7 Flash vs Claude Sonnet 5.

    Which model should you actually run in 2026?

    Match the model to the constraint that is actually binding you — license risk, serving budget, or raw capability. For most teams under 10 billion tokens a month, the answer is an API, and it probably is not Qwen’s.

    Your situationBest choiceWhy
    Under 5B tokens/monthQwen3.8-Max API$2/$6 with $0.25 cached input beats any cluster you can rent
    Coding agents at scaleDeepSeek V4 Pro80.6% SWE-bench, MIT license, $0.435/$0.87
    Air-gapped or regulated deploymentGLM-5.2744B/40B is the only one that fits comfortably on one node
    MaaS provider above $50M revenueAnything MIT-licensedQwen’s license forces a negotiation you will lose
    Research, agents, long-horizon tasksQwen3.8-Max APIOSWorld 86.1 and PaperBench 93.0 lead the field
    Above 20B tokens/month, MIT requiredSelf-host DeepSeek V4 Pro49B active params serve far cheaper than Qwen’s 95B

    Frequently asked questions about Qwen3.8-Max open weights

    Are the Qwen3.8-Max open weights free?

    Free to download and free to use commercially below $50 million in annual revenue. Above that threshold, Model-as-a-Service and AI assistant providers must negotiate a separate commercial license with Alibaba.

    Is the open checkpoint the same model as the Qwen3.8-Max API?

    No. The open weights are text-only and thinking-mode only, without the vision input and native 1M-token context that the hosted API provides.

    How much hardware do I need to run Qwen3.8-Max?

    The full-precision release is roughly 4.89 TB across 224 files, with an FP8 variant at about half that. Plan for multiple eight-GPU nodes, not one.

    Is Qwen3.8-Max better than DeepSeek V4 Pro?

    Not for coding. DeepSeek V4 Pro scores 80.6% on SWE-bench versus Qwen’s 67.7 on SWE-bench Pro, ships under MIT, and costs roughly a fifth as much per token.

    Why did Alibaba drop Apache 2.0?

    Commercial strategy. Apache 2.0 would have let rival inference providers build businesses on Alibaba’s research at zero cost. The revenue gate keeps distribution while capturing the upside at scale.

    When does self-hosting Qwen3.8-Max become cheaper than the API?

    Around 6.5 billion tokens a month in the best case, and realistically past 13 billion once you run two nodes. Add engineering headcount and the crossover pushes past 20 billion.

    Did the Qwen3.8-27B model ever ship?

    Not as of mid-August 2026. The smaller checkpoint was announced alongside Max but has not appeared on Hugging Face.

    The bottom line

    Use the Qwen3.8-Max API if you are under roughly five billion tokens a month and you need agentic or research performance — OSWorld 86.1 and PaperBench 93.0 are genuinely class-leading, and $2/$6 is fair for that tier.

    Do not self-host it. The break-even sits above 13 billion tokens a month at realistic node counts, and the model you would be hosting is the stripped text-only build, not the one that posts those benchmark numbers.

    If you are building anything you intend to sell inference on, pick DeepSeek V4 Pro or GLM-5.2 instead. MIT licensing costs nothing at $50 million in revenue. Alibaba’s license costs you a negotiation with a company that also competes with you.

    The open-weights headline was real. The gift was not. Qwen3.8-Max is the most capable model you can legally download this month and the one with the most expensive fine print — and for once, both halves of that sentence matter equally. For more on how speed and cost trade off at the frontier, see our analysis of what 14x inference speed actually costs.

    Sources

  • OpenAI Ultrafast vs Claude Fast Mode: What 14x Speed Actually Costs

    OpenAI Ultrafast vs Claude Fast Mode is not a close race on speed. OpenAI’s new mode runs GPT-5.6 Sol at up to 750 output tokens per second — 14x standard, on Cerebras silicon. Anthropic’s Fast mode delivers up to 2.5x for an exact 2x price premium. Anthropic publishes its price; OpenAI has not. That single gap decides who wins.

    Both landed on August 13, 2026. Both sell the same thing: the same model weights, running faster, for more money.

    The interesting question is not which is faster. It is what a second of latency is actually worth on your P&L.

    What is OpenAI Ultrafast mode?

    Ultrafast is a speed tier for GPT-5.6 Sol, not a new model. OpenAI’s announcement puts it at up to 14x standard processing and up to 750 output tokens per second, in limited preview for a small group of customers, expanding “as capacity grows.”

    OpenAI framed the pitch bluntly: “Until now, getting real-time speed typically meant choosing a smaller or more specialized model.”

    No price has been published. That omission is the whole story.

    The Cerebras hardware behind the number

    Ultrafast runs on Cerebras Wafer-Scale Engine chips. Per Cerebras’s own release, each wafer-sized chip carries 44 GB of on-chip SRAM, so model weights stay resident instead of shuttling to external memory.

    That architecture is why the multiplier is 14x and not 1.4x. It is also why capacity is rationed — wafer-scale supply does not scale like renting more GPUs.

    OpenAI Ultrafast vs Claude Fast Mode: how do the speed claims compare?

    Anthropic’s Fast mode delivers up to 2.5x higher output tokens per second on Claude Opus 5 and Opus 4.8, per Anthropic’s documentation. Cerebras claims Ultrafast is 5x faster than Opus 4.8 in Fast mode and 11x faster than Claude Fable 5. Treat competitor-run numbers with care.

    Speed tier Model Speed claim Input / 1M Output / 1M Premium
    OpenAI Ultrafast GPT-5.6 Sol Up to 14x; 750 tok/sec Not published Not published Undisclosed
    GPT-5.6 Sol (standard) GPT-5.6 Sol Baseline $2.50 $15.00
    Claude Fast mode Opus 5 / Opus 4.8 Up to 2.5x OTPS $10.00 $50.00 Exactly 2x
    Claude Opus 5 (standard) Opus 5 Baseline $5.00 $25.00
    Claude Fable 5 Fable 5 Standard speed $10.00 $50.00
    Sources: OpenAI Ultrafast preview, Cerebras press release, Anthropic pricing and Fast mode docs (August 2026).

    Reading the Cerebras claims honestly

    Cerebras also reports a 7x faster completion on Humanity’s Last Exam — 11-plus hours against 3-plus days — and a 5.6x end-to-end speedup on GDP-Val.

    Those are vendor numbers from the party selling the chips. But the direction is consistent with the architecture, and OpenAI’s own 750 tokens-per-second figure is published independently.

    If the 5x claim holds, Opus 4.8 in Fast mode lands near 150 output tokens per second. Fable 5 sits near 68. Both are derived, not published.

    How much does Claude Fast Mode actually cost?

    Exactly double. Opus 5 lists at $5/$25 per million tokens; Fast mode lists at $10/$50, per Anthropic’s pricing page. On an 80/20 input-output mix that is $18.00 per million blended against $9.00 standard.

    Here is the detail nobody flags: $10/$50 is also the exact list price of Claude Fable 5, Anthropic’s top tier.

    So Opus 5 at 2.5x speed costs precisely what Anthropic’s most capable model costs at normal speed. Speed and frontier intelligence are priced identically. That is a deliberate pricing choice, and it caps how much speed can ever be worth inside Anthropic’s own lineup.

    The hidden costs of Fast mode

    The sticker premium is not the full bill. Anthropic’s docs list several constraints that quietly raise effective cost:

    • Cache invalidation: switching between speeds clears cached prefixes. A fallback to standard speed is a guaranteed cache miss.
    • No Batch API: the 50% batch discount is unavailable in Fast mode.
    • No Priority Tier: incompatible with committed-capacity contracts.
    • API only: unavailable on Bedrock, Google Cloud, and Microsoft Foundry.
    • Separate rate limits: Fast mode has its own quota and returns 429s independently of standard Opus limits.
    • TTFT unchanged: only output throughput improves, so short responses barely benefit.

    Multipliers stack too. Prompt caching and US-only data residency apply on top of the $10/$50 base, not instead of it.

    What will OpenAI Ultrafast cost?

    OpenAI has not said. Neither the announcement, the Cerebras release, nor TechCrunch’s coverage carries a number. So model it: GPT-5.6 Sol lists at $2.50/$15.00, a $5.00 blended rate. Every plausible premium still lands under Anthropic.

    Scenario Input / 1M Output / 1M Blended (80/20) vs. Claude Fast mode
    Sol at standard price $2.50 $15.00 $5.00 72% cheaper
    Sol at Anthropic’s 2x premium $5.00 $30.00 $10.00 44% cheaper
    Sol at a 3x premium $7.50 $45.00 $15.00 17% cheaper
    Sol at a 3.6x premium $9.00 $54.00 $18.00 Parity
    Claude Opus 5 Fast mode $10.00 $50.00 $18.00
    Modeled from published GPT-5.6 Sol list pricing. OpenAI has not disclosed Ultrafast pricing.

    OpenAI would need to charge a 3.6x premium just to match Anthropic’s blended Fast mode rate — while delivering roughly 5x the throughput. That is the box Anthropic is now in.

    Is paying for faster inference worth it?

    Only when latency blocks something billable. Speed premiums pay for themselves in interactive and long-horizon agent work, and waste money everywhere else. The test is simple: if the output goes into a queue, you are burning margin on throughput nobody is waiting for.

    Run the arithmetic on a 10-million-output-token job — roughly a large agentic refactor or a bulk document pipeline.

    Configuration Output cost Throughput Wall-clock time
    GPT-5.6 Sol Ultrafast Price undisclosed 750 tok/sec ~3.7 hours
    Claude Opus 4.8 Fast mode $500 ~150 tok/sec (derived) ~18.5 hours
    Claude Fable 5 $500 ~68 tok/sec (derived) ~40.8 hours
    Claude Opus 5 standard $250 Baseline ~46 hours (derived)
    GPT-5.6 Sol standard $150 Baseline ~52 hours (derived)
    Costs from published list prices. Throughput for Claude tiers derived from Cerebras’s comparative claims, not vendor-published figures.

    The spread between $150 and $500 is real money, but it is not what decides this. A pipeline that clears in under four hours runs inside a working day. One that takes 46 hours does not.

    Which speed tier should you buy for which job?

    Match the tier to whether a human is waiting. Interactive products and incident response justify a premium; overnight batch work never does. OpenAI named the same set of use cases — incident response, fraud detection, real-time support, e-commerce assistance — which tells you where it expects the money to come from.

    Use case Best tier Why
    Real-time support and copilots OpenAI Ultrafast 750 tok/sec makes synchronous UX viable
    Incident response and on-call triage OpenAI Ultrafast Minutes of downtime cost more than tokens
    Long-horizon agent runs Ultrafast, or Opus 5 Fast mode 7x faster completion on long tasks, per Cerebras
    High-stakes reasoning, human in the loop Claude Opus 5 Fast mode 2.5x OTPS at a known, published price
    Overnight batch and bulk processing Standard tiers with Batch API Fast modes forfeit the 50% batch discount
    Short responses and classification Standard tiers Fast mode does not improve time to first token
    Bedrock, Vertex, or Foundry deployments Standard tiers only Claude Fast mode is first-party API only

    Who actually wins financially?

    Cerebras. The chipmaker went public on May 14, 2026, popping 68% on debut to a roughly $95 billion market cap, per CNBC. Powering OpenAI’s flagship speed tier converts that valuation from a thesis into a revenue line.

    The second winner is buyers with leverage. A priced 2.5x tier now sits next to an unpriced 14x tier, and Anthropic set the anchor first — the same defensive posture visible when it took a $2 trillion valuation and spent $6 billion on getting cheaper.

    The loser is anyone who assumed inference costs only fall. DeepSeek raised prices up to 1,100% overnight this week. Speed is being sold as a separate SKU, priced above the model itself. That is the opposite of commoditization — and it sits directly against the token-price collapse we tracked in Gemini 3.7 Flash versus Claude Sonnet 5.

    Frequently asked questions

    How fast is OpenAI Ultrafast mode?

    Up to 14x standard processing and up to 750 output tokens per second on GPT-5.6 Sol, running on Cerebras Wafer-Scale Engine hardware. It is in limited preview for a small group of customers.

    How much does OpenAI Ultrafast cost?

    OpenAI has not published pricing. GPT-5.6 Sol lists at $2.50 input and $15.00 output per million tokens at standard speed, so any premium starts from there.

    How much does Claude Fast Mode cost?

    $10 input and $50 output per million tokens for Claude Opus 5 and Opus 4.8 — exactly double the standard $5/$25. That is $18.00 blended on an 80/20 mix.

    Does Claude Fast Mode work with the Batch API?

    No. Fast mode is incompatible with the Batch API, Priority Tier, and partner clouds including Bedrock, Google Cloud, and Microsoft Foundry. It is first-party Claude API only.

    Does Fast mode make responses start faster?

    No. Anthropic states the benefit is output tokens per second, not time to first token. Short responses see little improvement.

    Is Ultrafast a different model from GPT-5.6 Sol?

    No. Both Ultrafast and Claude Fast mode run identical model weights at higher throughput. Capability does not change; only speed and price do.

    Can I get access to Ultrafast today?

    Only through the limited preview. OpenAI and Cerebras both direct interested customers to registration forms, with expansion tied to available wafer-scale capacity.

    The bottom line

    If you can get into the Ultrafast preview, take it. A 14x throughput tier at 750 tokens per second changes what an agent can finish inside a working day, and OpenAI would have to charge a 3.6x premium over Sol’s list price before it even reaches Anthropic’s blended Fast mode rate.

    Buy Claude Opus 5 Fast mode when you need Anthropic’s reasoning and a price you can put in a budget today. Known cost beats unknown cost when finance has to sign.

    Buy neither for anything queued. Batch and standard tiers are 50% cheaper still, and Fast mode explicitly forfeits that discount. The decisive variable is whether a person — or a paying customer — is waiting on the tokens. If nobody is, every dollar of speed premium is waste.

    Sources

  • Gemini 3.7 Flash vs Claude Sonnet 5: Which Wins on Cost Per Coding Point?

    Gemini 3.7 Flash vs Claude Sonnet 5 comes down to one number: cost per benchmark point. Google’s new workhorse matches Sonnet 5 on production coding evals while listing at $0.75/$3.75 per million tokens against Anthropic’s $2/$10. That is roughly 2.7x cheaper for equal-or-better coding output. Sonnet 5 still wins on the hardest reasoning tests. For agent workloads that burn tokens all day, Flash wins on money.

    Google shipped Gemini 3.7 Flash on August 13, 2026 — three weeks after Gemini 3.6 Flash. The release matters less as a launch and more as a repricing event. When a cheap model closes the coding gap with a premium model, every AI budget line gets renegotiated.

    We priced all three frontier options against their published benchmarks using vendor list prices. The result is not close.

    How much does Gemini 3.7 Flash cost compared to Claude Sonnet 5?

    Gemini 3.7 Flash lists at $0.75 per million input tokens and $3.75 per million output through December 31, 2026, per Google’s official Gemini API pricing page. Claude Sonnet 5 lists at $2.00 and $10.00. On an 80/20 input-output mix, that is $1.35 versus $3.60 per million tokens.

    Anthropic also settled a question that had been hanging over Sonnet 5’s price. Its pricing documentation now states that the introductory $2/$10 rate is permanent and the scheduled September 1, 2026 increase to $3/$15 “will not occur.”

    That was a defensive move. It did not close the gap.

    The full price and spec comparison

    Model Input / 1M Output / 1M Blended (80/20) Context Batch in / out
    Gemini 3.7 Flash $0.75 $3.75 $1.35 1M in / 64K out $0.375 / $1.875
    Gemini 3.7 Flash (from Jan 1, 2027) $1.50 $7.50 $2.70 1M in / 64K out $0.75 / $3.75
    Claude Sonnet 5 $2.00 $10.00 $3.60 1M at standard rate $1.00 / $5.00
    GPT-5.6 Terra $1.00 $6.00 $2.00 Long-context tier priced separately
    GPT-5.6 Sol $2.50 $15.00 $5.00 Long-context tier priced separately
    Claude Opus 5 $5.00 $25.00 $9.00 1M at standard rate $2.50 / $12.50
    Sources: Google Gemini API pricing, Anthropic pricing docs, OpenAI API pricing (August 2026).

    The January 1, 2027 price cliff

    Google’s $0.75 rate is introductory. On January 1, 2027 it doubles to $1.50/$7.50, which lifts the blended cost to $2.70.

    Even then, Flash stays 25% under Sonnet 5. But the deepest discount window is four and a half months wide, and it is the single best arbitrage on the table right now.

    The tokenizer tax nobody prices in

    Anthropic’s documentation carries a note most comparison tables ignore: Claude 4.7 and later models use a newer tokenizer that “produces approximately 30% more tokens for the same text.”

    Sticker price is per token. Your bill is per document. If that 30% applies to your workload, Sonnet 5’s effective cost per page of English moves closer to $4.70 blended — over 3x Flash’s introductory rate.

    Gemini 3.7 Flash vs Claude Sonnet 5: which is better for coding?

    Flash wins on production coding and agentic execution; Sonnet 5 wins on long-horizon reasoning. On Google’s published model-card comparisons, Flash takes FrontierCode 1.1 at 43.6% against Sonnet 5’s 42.7%, and crushes it on AutomationBench, 30.4% to 10.7%. Sonnet 5 answers on GDPval and Agent’s Last Exam.

    The benchmark table below is vendor-stated from Google’s model card, tabulated independently by DataCamp and other outlets. Treat first-party numbers with the usual skepticism — but they are consistent across sources.

    Benchmark Gemini 3.7 Flash Claude Sonnet 5 GPT-5.6 Terra
    FrontierCode 1.1 (production code) 43.6% 42.7% 41.3%
    DeepSWE v1.1 (long-horizon SWE) 65.3% 53.8% 69.6%
    Terminal-bench 2.1 85.8% 80.4% 87.4%
    WebDev Arena (Elo) 1588 1541 1523
    AutomationBench 30.4% 10.7% 23.6%
    GDPval-AA v2 (Elo) 1525 1598 1578
    GDM-MRCR v2, 128k (recall) 97.0% 81.5% 93.5%
    Agent’s Last Exam 26.3% 33.3% 28.0%
    Vendor-stated scores from Google’s Gemini 3.7 Flash model card, August 13, 2026.

    Cost per benchmark point: the number that decides it

    Divide blended cost by score and the argument ends.

    • FrontierCode 1.1: Flash costs $0.031 per point. Sonnet 5 costs $0.084. Terra costs $0.048.
    • DeepSWE v1.1: Flash $0.021 per point, Terra $0.029, Sonnet 5 $0.067.
    • AutomationBench: Flash delivers 2.8x Sonnet 5’s score at 37% of the price.
    • Long-context recall (MRCR 128k): Flash leads by 15.5 points and costs 63% less.

    Sonnet 5 is charging a 2.7x premium to lose a coding benchmark by 0.9 points. That is not a defensible position in a procurement meeting.

    Where Claude Sonnet 5 still earns its price

    Two places. Sonnet 5 leads GDPval-AA v2 at 1598 Elo against Flash’s 1525 — that benchmark tracks economically valuable knowledge work, not code. It also leads Agent’s Last Exam, 33.3% to 26.3%.

    If your workload is legal analysis, financial modeling, or research synthesis rather than shipping code, the premium is arguable. If it is code, it is not.

    How does GPT-5.6 Terra change the math?

    Terra is the quiet value play, and most launch-day comparison tables mispriced it. OpenAI’s official pricing page lists gpt-5.6-terra at $1.00 input and $6.00 output per million, with a separate long-context tier at $2.00/$9.00 — not the $2.00/$12.00 figure that circulated all week.

    At the correct list price, Terra blends to $2.00 per million. It also posts the best DeepSWE v1.1 score in the group at 69.6% and the best Terminal-bench 2.1 at 87.4%.

    Terra’s cached input runs $0.10 per million, half of Anthropic’s $0.20 cache-hit rate for Sonnet 5. For retrieval-heavy agents replaying the same system prompt thousands of times a day, that difference compounds fast.

    What does this actually cost at production volume?

    Take a mid-size agent workload: 500 million input tokens and 100 million output tokens per month. That is a realistic footprint for a coding assistant serving a 50-engineer team. The spread between the cheapest and most expensive option is $4,250 a month.

    Model Monthly cost Annualized vs. Sonnet 5
    Gemini 3.7 Flash (intro) $750 $9,000 −$15,000/yr
    GPT-5.6 Terra $1,100 $13,200 −$10,800/yr
    Gemini 3.7 Flash (2027 rate) $1,500 $18,000 −$6,000/yr
    Claude Sonnet 5 $2,000 $24,000
    GPT-5.6 Sol $2,750 $33,000 +$9,000/yr
    Claude Opus 5 $5,000 $60,000 +$36,000/yr
    Calculated from vendor list prices at 500M input / 100M output tokens per month.

    Batch processing cuts all of it roughly in half. Gemini 3.7 Flash drops to $0.375/$1.875 through year-end; Sonnet 5 drops to $1.00/$5.00. The ranking does not change.

    Which model should you use for which job?

    Match the model to the failure mode you can least afford. Coding agents that run unsupervised for hours need throughput and cheap retries. Client-facing analysis needs reasoning depth. Nothing here is a universal answer, and paying Opus prices for autocomplete is how AI budgets die.

    Use case Best choice Why
    High-volume coding agents Gemini 3.7 Flash Top FrontierCode score at 37% of Sonnet 5’s blended price
    Long-horizon autonomous SWE GPT-5.6 Terra Leads DeepSWE (69.6%) and Terminal-bench 2.1 (87.4%)
    Front-end and web generation Gemini 3.7 Flash WebDev Arena Elo 1588, ahead of both rivals
    Legal, financial, research synthesis Claude Sonnet 5 Top GDPval-AA v2 Elo at 1598
    Million-token document pipelines Gemini 3.7 Flash 97.0% MRCR recall at 128k, cheapest per token
    Hardest reasoning, cost no object Claude Opus 5 Frontier tier — $9.00 blended, use sparingly
    Bulk offline processing Gemini 3.7 Flash (batch) $0.375 / $1.875 through Dec 31, 2026

    Is switching to Gemini 3.7 Flash worth it in 2026?

    Yes, if your token spend clears roughly $1,000 a month. Below that, migration engineering costs more than it saves. Above it, the savings compound — and Google’s three-week release cadence means the model you migrate to keeps improving without a renegotiation.

    The strategic read is bigger than one model. Frontier-tier coding capability is commoditizing on a quarterly clock, and price is the only lever customers can still feel. We saw the other side of that trade this week when DeepSeek raised prices by up to 1,100% overnight — the cheap-inference era is being rationed, not extended.

    Anthropic is spending to stay in the fight. It reached a $2 trillion valuation and put $6 billion into getting cheaper. Cancelling the September price increase is the visible half of that strategy.

    One caution before you point an agent at production: capability and autonomy scale together. The same agentic execution that makes Flash cheap to run is what let an AI agent crack 85 accounts in four days. Sandbox accordingly.

    Frequently asked questions

    Is Gemini 3.7 Flash actually cheaper than Claude Sonnet 5?

    Yes. $0.75/$3.75 per million tokens versus $2.00/$10.00 — about 2.7x cheaper on an 80/20 blend. The introductory rate holds through December 31, 2026, then doubles to $1.50/$7.50.

    Does Gemini 3.7 Flash beat Claude Sonnet 5 at coding?

    On Google’s published card, yes — narrowly on FrontierCode 1.1 (43.6% vs 42.7%), decisively on DeepSWE v1.1 (65.3% vs 53.8%) and AutomationBench (30.4% vs 10.7%). GPT-5.6 Terra still leads DeepSWE overall at 69.6%.

    What is Gemini 3.7 Flash’s context window?

    One million input tokens and up to 65,536 output tokens, with a March 2026 knowledge cutoff. Claude Sonnet 5 also offers a 1M-token window at standard per-token pricing.

    Did Claude Sonnet 5’s price go up on September 1, 2026?

    No. Anthropic’s documentation confirms the scheduled increase to $3/$15 per million tokens will not occur. The $2/$10 introductory rate is now the standard price.

    How much does GPT-5.6 Terra cost?

    $1.00 input and $6.00 output per million tokens on the standard tier, with a long-context tier at $2.00/$9.00. Cached input is $0.10 per million.

    Where can I use Gemini 3.7 Flash today?

    Google AI Studio, Android Studio, Google Antigravity, the Gemini Enterprise Agent Platform, and Gemini Spark for AI Pro and Ultra subscribers across 160+ countries.

    Should I run one model or mix them?

    Mix. Route bulk coding and document work to Flash, long-horizon autonomous tasks to Terra, and high-stakes analysis to Sonnet 5. Routing the majority of calls to the cheapest capable tier is where the savings actually come from.

    The bottom line

    Default to Gemini 3.7 Flash for coding and agent workloads. It wins or ties on the coding benchmarks that map to shipped software, costs $1.35 blended against Sonnet 5’s $3.60, and saves a 50-engineer team roughly $15,000 a year at the volumes above.

    Keep Claude Sonnet 5 for the narrow band where it leads: GDPval-style knowledge work and Agent’s Last Exam reasoning. Keep GPT-5.6 Terra for long-horizon autonomous engineering, where its 69.6% DeepSWE score is worth the extra $0.65 per million blended.

    And put a calendar reminder on December 31, 2026. That is when Google’s discount ends and this entire calculation gets re-run. The labs are shipping every three weeks now — the money chasing this market guarantees the next repricing is already in the pipeline.

    Sources