Tag: Cerebras

  • Cerebras vs Groq: Which Fast Inference API Is Worth the Money

    Cerebras vs Groq comes down to one trade. Cerebras serves GPT-OSS-120B at 1,641 tokens per second for $0.75 per million output tokens. Groq serves the same model at roughly 500 for $0.60. You pay about 25% more on output for roughly three times the speed. Buy Cerebras when a human or an agent is waiting. Buy Groq for batch work and overnight jobs.

    That trade just got sharper. On August 18, 2026, Cerebras announced the CS-4, a system it claims runs GPT-OSS-120B at more than 4,400 tokens per second per user.

    If that number survives contact with production traffic, the speed gap stops being a nice-to-have and starts being a product feature you can charge for.

    What changed in the Cerebras vs Groq race this week?

    Cerebras shipped a new generation of silicon and Nvidia now owns its main rival’s technology. Those two facts reshape the fast-inference market. The CS-4 raises Cerebras’ ceiling; the Nvidia-Groq deal means Groq’s aggressive pricing is no longer set by a scrappy independent.

    The CS-4 numbers that matter

    Per the Cerebras announcement, the CS-4 delivers 750 PFLOPS of AI compute against the CS-3’s 125 PFLOPS. Memory bandwidth jumps from 21.6 to 129.6 petabytes per second.

    The WSE-3 Turbo processor behind it packs 4 trillion transistors and 900,000 AI cores across 46,225 square millimeters of silicon, with 44GB of on-chip SRAM.

    Wafer-to-wafer latency drops from 5 microseconds to 2. Cerebras also claims up to 10x the throughput per watt versus the CS-3, and support for models above 50 trillion parameters.

    The CS-4 product page adds a second claim worth watching: more than 1,000 tokens per second on models exceeding 10 trillion parameters. First shipments began in Q3 2026.

    CTO Sean Lie framed the pitch in agent terms: “Being 30 times faster gives an agentic system room for significantly more reasoning, verification, or tool use in the same wall-clock time.”

    Why Nvidia now sits on both sides

    Groq is no longer an independent challenger. CNBC reported on December 24, 2025 that Nvidia agreed to buy Groq’s assets for about $20 billion — its largest deal on record.

    The Groq API still runs and still undercuts Cerebras. But the pricing that made Groq attractive is now a line item inside the company that also sells the GPUs Groq was built to beat.

    That matters for anyone building a business on a specific cost per token. Cerebras is the last large pure-play fast-inference vendor with its own silicon and its own incentive to keep prices down.

    How much does fast AI inference cost per million tokens?

    Cerebras is the most expensive way to run GPT-OSS-120B among mainstream providers. Groq sits mid-pack. Commodity GPU serverless tiers cost a fraction of both. The spread on the identical open-weights model is roughly 7x on output tokens, which is far wider than most teams assume.

    Here is the pricing snapshot for GPT-OSS-120B as tracked by PricePerToken on August 21, 2026, with measured speed where it is published.

    Provider Input $/M Output $/M 1M in + 1M out Measured output speed
    Cerebras $0.35 $0.75 $1.10 1,641 tok/s
    Groq $0.15 $0.60 $0.75 ~500 tok/s
    SambaNova $0.14 $0.95 $1.09 Not published
    Together AI $0.15 $0.60 $0.75 Not published
    Amazon Bedrock $0.15 $0.60 $0.75 Not published
    Baseten $0.10 $0.50 $0.60 Not published
    Google $0.09 $0.36 $0.45 Not published
    Fireworks $0.10 $0.10 $0.20 Not published
    DeepInfra $0.037 $0.170 $0.207 Not published
    OpenRouter $0.030 $0.170 $0.200 Not published
    Cerebras speed from Artificial Analysis; Groq speed from CloudZero, May 2026.

    Run a balanced million-in, million-out workload and Cerebras costs $1.10 against Groq’s $0.75 — a 47% premium. On output tokens alone the gap narrows to 25%.

    Both are expensive next to Fireworks at $0.20 or the OpenRouter route at $0.20 for the same weights.

    The discount lever most teams forget

    Groq’s list price is not its real price. CloudZero notes that Groq’s Batch API and prompt caching each cut rates by 50%, and the two stack to roughly 25% of on-demand pricing.

    Applied to GPT-OSS-120B, that pushes effective output cost toward $0.15 per million. Cerebras publishes no equivalent public stacking discount; its pricing page lists a $5 free trial, a $10 self-serve developer tier, and custom enterprise rates.

    For any workload that tolerates a batch window, Groq is not 20% cheaper. It is closer to 5x cheaper.

    Which is faster in practice, Cerebras or Groq?

    Cerebras wins today, and not narrowly. Independent measurement puts Cerebras at 1,641 tokens per second on GPT-OSS-120B with 0.46 seconds to first token. Groq’s published figure for the same model is around 500 tokens per second. That is a 3.3x throughput advantage before the CS-4 ships at volume.

    Artificial Analysis also clocks Cerebras at 1,402 tokens per second on Gemma 4 31B in reasoning mode, at a blended $0.24 per million.

    What the speed gap costs in real money

    Take a 100,000-token generation — a long agent trace or a full document rewrite.

    • Cerebras today: 61 seconds at 1,641 tok/s, costing $0.075 in output tokens.
    • Groq today: 200 seconds at 500 tok/s, costing $0.060.
    • Cerebras CS-4 claim: 23 seconds at 4,400 tok/s.
    • The math: 1.5 cents buys back 139 seconds — roughly 11 cents per minute of latency removed.

    Eleven cents a minute is trivial if a customer is watching a cursor blink. It is indefensible if the job runs at 3 a.m. and nobody reads the output until morning.

    One honest caveat: the 4,400 tok/s figure is a vendor claim tied to hardware that only started shipping this quarter. The 1,641 figure is measured on the live API. Do not budget against the former.

    Is Cerebras worth the premium for AI agents?

    Yes, for interactive and multi-step agent work — and this is the only case where the premium clearly pays. Agent loops multiply latency: ten sequential tool calls at 200 seconds each is a 33-minute task. The same loop at Cerebras speed finishes in about 10 minutes.

    That compounding is the whole argument. A single completion at 500 tokens per second feels fine. Twenty of them chained behind a task does not.

    Who should buy what

    Use case Pick Why
    Live chat, copilots, voice Cerebras 0.46s TTFT and 1,641 tok/s; latency is the product
    Multi-step agents with tool calls Cerebras Per-step latency compounds across the loop
    Overnight batch, evals, data labeling Groq (Batch API) Stacked discounts reach ~25% of list
    High-volume, cost-capped production Fireworks / DeepInfra $0.20 per 1M in + 1M out on identical weights
    Closed frontier models Neither Both serve open weights only
    Sustained 24/7 single-model load Self-host Fixed GPU cost beats per-token above a break-even

    That last row matters more as open weights close the quality gap. We ran the self-hosting break-even in our Qwen3.8-Max open weights breakdown, and the logic holds here.

    Note the hard limit on both vendors: neither serves closed frontier models. If your stack depends on the newest proprietary coder, this comparison does not apply — see our GLM-5.3 vs DeepSeek V4 Pro comparison for the open-weight options that do run here.

    What do the financials say about who wins?

    Cerebras has the better technology and the shakier income statement. It went public on May 14, 2026, raising $5.5 billion. The stock priced at $185, opened at $385 for a 108% pop, and closed at $311 — a $66 billion valuation. Three months later the market repriced it hard.

    The Q2 miss

    On August 12, 2026, Cerebras reported Q2 revenue of $180.1 million against a $193.6 million consensus, per Investing.com. Adjusted EPS came in at a $2.98 loss versus an expected $0.18 loss. Shares fell 14% after hours.

    Core revenue still grew 103% year over year to $209.9 million, and the company guided FY2026 core revenue to $880–890 million. Growth is not the problem. Margin is: guided core operating margin sits at negative 19% to negative 17%.

    Why that should affect your buying decision

    A vendor losing money on every wafer has two exits: raise prices or get acquired. Cerebras’ 2025 revenue was $510 million on 76% growth with $237.8 million of net income, so the balance sheet is not fragile — but the 2026 trajectory is being funded, not earned.

    TrendForce values the company’s three-year OpenAI partnership at over $20 billion. That is concentration risk dressed as a moat.

    The broader point TrendForce makes is the one to internalize: inference is a recurring cost tied directly to revenue, while training is a one-time R&D expense. Its example is brutal — Taalas’ HC1 delivers Llama 3.1 8B at 0.75 cents per million tokens against 3.79 cents on an Nvidia B200.

    Specialized silicon is roughly five times cheaper per token than general-purpose GPUs at that scale. That is why the price you lock in today is unlikely to be the price in twelve months.

    Frequently asked questions about Cerebras vs Groq

    Is Cerebras faster than Groq?

    Yes. Artificial Analysis measures Cerebras at 1,641 tokens per second on GPT-OSS-120B; Groq’s published figure for the same model is around 500. Cerebras also leads on time to first token at 0.46 seconds.

    Is Groq cheaper than Cerebras?

    Yes. On GPT-OSS-120B, Groq charges $0.15 input and $0.60 output per million tokens versus Cerebras at $0.35 and $0.75. With Groq’s stacked batch and cache discounts, the effective gap widens sharply.

    Does Nvidia own Groq now?

    Nvidia agreed in December 2025 to acquire Groq’s assets for about $20 billion, CNBC reported. The Groq API continues to operate and continues to publish its own pricing.

    Can I run Claude or GPT-5 on Cerebras or Groq?

    No. Both providers serve open-weights models only — GPT-OSS, Llama, Qwen, Gemma and similar. Closed frontier models stay on their vendors’ own APIs.

    When does the CS-4 speed actually arrive for API users?

    Cerebras says first CS-4 shipments began in Q3 2026. The 4,400 tokens per second per user figure is a vendor claim on new hardware, not yet an independently measured API result.

    Is paying for faster inference ever worth it?

    Only when latency is visible to a customer or compounds across an agent loop. At roughly 11 cents per minute of latency removed, speed is cheap for interactive products and pure waste for background jobs. We covered the same trade at the frontier-model layer in OpenAI Ultrafast vs Claude Fast Mode.

    The bottom line

    Route by whether something is waiting. If a human or an agent loop blocks on the token stream, Cerebras is worth its 25% output premium — 3.3x measured throughput for that price is one of the better deals in AI infrastructure, and the CS-4 should widen it.

    If nothing is waiting, Cerebras is a rounding-error upgrade you are overpaying for. Send batch and evaluation traffic to Groq’s Batch API at roughly a quarter of list, or to Fireworks and DeepInfra at $0.20 per million in and out.

    The strategic read is less comfortable. Nvidia owns Groq’s technology, Cerebras is losing money at negative 17% to 19% core operating margin, and specialized silicon is already showing five-fold cost advantages per token. Sign nothing longer than twelve months.

    Sources

  • OpenAI Ultrafast vs Claude Fast Mode: What 14x Speed Actually Costs

    OpenAI Ultrafast vs Claude Fast Mode is not a close race on speed. OpenAI’s new mode runs GPT-5.6 Sol at up to 750 output tokens per second — 14x standard, on Cerebras silicon. Anthropic’s Fast mode delivers up to 2.5x for an exact 2x price premium. Anthropic publishes its price; OpenAI has not. That single gap decides who wins.

    Both landed on August 13, 2026. Both sell the same thing: the same model weights, running faster, for more money.

    The interesting question is not which is faster. It is what a second of latency is actually worth on your P&L.

    What is OpenAI Ultrafast mode?

    Ultrafast is a speed tier for GPT-5.6 Sol, not a new model. OpenAI’s announcement puts it at up to 14x standard processing and up to 750 output tokens per second, in limited preview for a small group of customers, expanding “as capacity grows.”

    OpenAI framed the pitch bluntly: “Until now, getting real-time speed typically meant choosing a smaller or more specialized model.”

    No price has been published. That omission is the whole story.

    The Cerebras hardware behind the number

    Ultrafast runs on Cerebras Wafer-Scale Engine chips. Per Cerebras’s own release, each wafer-sized chip carries 44 GB of on-chip SRAM, so model weights stay resident instead of shuttling to external memory.

    That architecture is why the multiplier is 14x and not 1.4x. It is also why capacity is rationed — wafer-scale supply does not scale like renting more GPUs.

    OpenAI Ultrafast vs Claude Fast Mode: how do the speed claims compare?

    Anthropic’s Fast mode delivers up to 2.5x higher output tokens per second on Claude Opus 5 and Opus 4.8, per Anthropic’s documentation. Cerebras claims Ultrafast is 5x faster than Opus 4.8 in Fast mode and 11x faster than Claude Fable 5. Treat competitor-run numbers with care.

    Speed tier Model Speed claim Input / 1M Output / 1M Premium
    OpenAI Ultrafast GPT-5.6 Sol Up to 14x; 750 tok/sec Not published Not published Undisclosed
    GPT-5.6 Sol (standard) GPT-5.6 Sol Baseline $2.50 $15.00
    Claude Fast mode Opus 5 / Opus 4.8 Up to 2.5x OTPS $10.00 $50.00 Exactly 2x
    Claude Opus 5 (standard) Opus 5 Baseline $5.00 $25.00
    Claude Fable 5 Fable 5 Standard speed $10.00 $50.00
    Sources: OpenAI Ultrafast preview, Cerebras press release, Anthropic pricing and Fast mode docs (August 2026).

    Reading the Cerebras claims honestly

    Cerebras also reports a 7x faster completion on Humanity’s Last Exam — 11-plus hours against 3-plus days — and a 5.6x end-to-end speedup on GDP-Val.

    Those are vendor numbers from the party selling the chips. But the direction is consistent with the architecture, and OpenAI’s own 750 tokens-per-second figure is published independently.

    If the 5x claim holds, Opus 4.8 in Fast mode lands near 150 output tokens per second. Fable 5 sits near 68. Both are derived, not published.

    How much does Claude Fast Mode actually cost?

    Exactly double. Opus 5 lists at $5/$25 per million tokens; Fast mode lists at $10/$50, per Anthropic’s pricing page. On an 80/20 input-output mix that is $18.00 per million blended against $9.00 standard.

    Here is the detail nobody flags: $10/$50 is also the exact list price of Claude Fable 5, Anthropic’s top tier.

    So Opus 5 at 2.5x speed costs precisely what Anthropic’s most capable model costs at normal speed. Speed and frontier intelligence are priced identically. That is a deliberate pricing choice, and it caps how much speed can ever be worth inside Anthropic’s own lineup.

    The hidden costs of Fast mode

    The sticker premium is not the full bill. Anthropic’s docs list several constraints that quietly raise effective cost:

    • Cache invalidation: switching between speeds clears cached prefixes. A fallback to standard speed is a guaranteed cache miss.
    • No Batch API: the 50% batch discount is unavailable in Fast mode.
    • No Priority Tier: incompatible with committed-capacity contracts.
    • API only: unavailable on Bedrock, Google Cloud, and Microsoft Foundry.
    • Separate rate limits: Fast mode has its own quota and returns 429s independently of standard Opus limits.
    • TTFT unchanged: only output throughput improves, so short responses barely benefit.

    Multipliers stack too. Prompt caching and US-only data residency apply on top of the $10/$50 base, not instead of it.

    What will OpenAI Ultrafast cost?

    OpenAI has not said. Neither the announcement, the Cerebras release, nor TechCrunch’s coverage carries a number. So model it: GPT-5.6 Sol lists at $2.50/$15.00, a $5.00 blended rate. Every plausible premium still lands under Anthropic.

    Scenario Input / 1M Output / 1M Blended (80/20) vs. Claude Fast mode
    Sol at standard price $2.50 $15.00 $5.00 72% cheaper
    Sol at Anthropic’s 2x premium $5.00 $30.00 $10.00 44% cheaper
    Sol at a 3x premium $7.50 $45.00 $15.00 17% cheaper
    Sol at a 3.6x premium $9.00 $54.00 $18.00 Parity
    Claude Opus 5 Fast mode $10.00 $50.00 $18.00
    Modeled from published GPT-5.6 Sol list pricing. OpenAI has not disclosed Ultrafast pricing.

    OpenAI would need to charge a 3.6x premium just to match Anthropic’s blended Fast mode rate — while delivering roughly 5x the throughput. That is the box Anthropic is now in.

    Is paying for faster inference worth it?

    Only when latency blocks something billable. Speed premiums pay for themselves in interactive and long-horizon agent work, and waste money everywhere else. The test is simple: if the output goes into a queue, you are burning margin on throughput nobody is waiting for.

    Run the arithmetic on a 10-million-output-token job — roughly a large agentic refactor or a bulk document pipeline.

    Configuration Output cost Throughput Wall-clock time
    GPT-5.6 Sol Ultrafast Price undisclosed 750 tok/sec ~3.7 hours
    Claude Opus 4.8 Fast mode $500 ~150 tok/sec (derived) ~18.5 hours
    Claude Fable 5 $500 ~68 tok/sec (derived) ~40.8 hours
    Claude Opus 5 standard $250 Baseline ~46 hours (derived)
    GPT-5.6 Sol standard $150 Baseline ~52 hours (derived)
    Costs from published list prices. Throughput for Claude tiers derived from Cerebras’s comparative claims, not vendor-published figures.

    The spread between $150 and $500 is real money, but it is not what decides this. A pipeline that clears in under four hours runs inside a working day. One that takes 46 hours does not.

    Which speed tier should you buy for which job?

    Match the tier to whether a human is waiting. Interactive products and incident response justify a premium; overnight batch work never does. OpenAI named the same set of use cases — incident response, fraud detection, real-time support, e-commerce assistance — which tells you where it expects the money to come from.

    Use case Best tier Why
    Real-time support and copilots OpenAI Ultrafast 750 tok/sec makes synchronous UX viable
    Incident response and on-call triage OpenAI Ultrafast Minutes of downtime cost more than tokens
    Long-horizon agent runs Ultrafast, or Opus 5 Fast mode 7x faster completion on long tasks, per Cerebras
    High-stakes reasoning, human in the loop Claude Opus 5 Fast mode 2.5x OTPS at a known, published price
    Overnight batch and bulk processing Standard tiers with Batch API Fast modes forfeit the 50% batch discount
    Short responses and classification Standard tiers Fast mode does not improve time to first token
    Bedrock, Vertex, or Foundry deployments Standard tiers only Claude Fast mode is first-party API only

    Who actually wins financially?

    Cerebras. The chipmaker went public on May 14, 2026, popping 68% on debut to a roughly $95 billion market cap, per CNBC. Powering OpenAI’s flagship speed tier converts that valuation from a thesis into a revenue line.

    The second winner is buyers with leverage. A priced 2.5x tier now sits next to an unpriced 14x tier, and Anthropic set the anchor first — the same defensive posture visible when it took a $2 trillion valuation and spent $6 billion on getting cheaper.

    The loser is anyone who assumed inference costs only fall. DeepSeek raised prices up to 1,100% overnight this week. Speed is being sold as a separate SKU, priced above the model itself. That is the opposite of commoditization — and it sits directly against the token-price collapse we tracked in Gemini 3.7 Flash versus Claude Sonnet 5.

    Frequently asked questions

    How fast is OpenAI Ultrafast mode?

    Up to 14x standard processing and up to 750 output tokens per second on GPT-5.6 Sol, running on Cerebras Wafer-Scale Engine hardware. It is in limited preview for a small group of customers.

    How much does OpenAI Ultrafast cost?

    OpenAI has not published pricing. GPT-5.6 Sol lists at $2.50 input and $15.00 output per million tokens at standard speed, so any premium starts from there.

    How much does Claude Fast Mode cost?

    $10 input and $50 output per million tokens for Claude Opus 5 and Opus 4.8 — exactly double the standard $5/$25. That is $18.00 blended on an 80/20 mix.

    Does Claude Fast Mode work with the Batch API?

    No. Fast mode is incompatible with the Batch API, Priority Tier, and partner clouds including Bedrock, Google Cloud, and Microsoft Foundry. It is first-party Claude API only.

    Does Fast mode make responses start faster?

    No. Anthropic states the benefit is output tokens per second, not time to first token. Short responses see little improvement.

    Is Ultrafast a different model from GPT-5.6 Sol?

    No. Both Ultrafast and Claude Fast mode run identical model weights at higher throughput. Capability does not change; only speed and price do.

    Can I get access to Ultrafast today?

    Only through the limited preview. OpenAI and Cerebras both direct interested customers to registration forms, with expansion tied to available wafer-scale capacity.

    The bottom line

    If you can get into the Ultrafast preview, take it. A 14x throughput tier at 750 tokens per second changes what an agent can finish inside a working day, and OpenAI would have to charge a 3.6x premium over Sol’s list price before it even reaches Anthropic’s blended Fast mode rate.

    Buy Claude Opus 5 Fast mode when you need Anthropic’s reasoning and a price you can put in a budget today. Known cost beats unknown cost when finance has to sign.

    Buy neither for anything queued. Batch and standard tiers are 50% cheaper still, and Fast mode explicitly forfeits that discount. The decisive variable is whether a person — or a paying customer — is waiting on the tokens. If nobody is, every dollar of speed premium is waste.

    Sources