Cerebras vs Groq: Which Fast Inference API Is Worth the Money

Cerebras vs Groq comes down to one trade. Cerebras serves GPT-OSS-120B at 1,641 tokens per second for $0.75 per million output tokens. Groq serves the same model at roughly 500 for $0.60. You pay about 25% more on output for roughly three times the speed. Buy Cerebras when a human or an agent is waiting. Buy Groq for batch work and overnight jobs.

That trade just got sharper. On August 18, 2026, Cerebras announced the CS-4, a system it claims runs GPT-OSS-120B at more than 4,400 tokens per second per user.

If that number survives contact with production traffic, the speed gap stops being a nice-to-have and starts being a product feature you can charge for.

What changed in the Cerebras vs Groq race this week?

Cerebras shipped a new generation of silicon and Nvidia now owns its main rival’s technology. Those two facts reshape the fast-inference market. The CS-4 raises Cerebras’ ceiling; the Nvidia-Groq deal means Groq’s aggressive pricing is no longer set by a scrappy independent.

The CS-4 numbers that matter

Per the Cerebras announcement, the CS-4 delivers 750 PFLOPS of AI compute against the CS-3’s 125 PFLOPS. Memory bandwidth jumps from 21.6 to 129.6 petabytes per second.

The WSE-3 Turbo processor behind it packs 4 trillion transistors and 900,000 AI cores across 46,225 square millimeters of silicon, with 44GB of on-chip SRAM.

Wafer-to-wafer latency drops from 5 microseconds to 2. Cerebras also claims up to 10x the throughput per watt versus the CS-3, and support for models above 50 trillion parameters.

The CS-4 product page adds a second claim worth watching: more than 1,000 tokens per second on models exceeding 10 trillion parameters. First shipments began in Q3 2026.

CTO Sean Lie framed the pitch in agent terms: “Being 30 times faster gives an agentic system room for significantly more reasoning, verification, or tool use in the same wall-clock time.”

Why Nvidia now sits on both sides

Groq is no longer an independent challenger. CNBC reported on December 24, 2025 that Nvidia agreed to buy Groq’s assets for about $20 billion — its largest deal on record.

The Groq API still runs and still undercuts Cerebras. But the pricing that made Groq attractive is now a line item inside the company that also sells the GPUs Groq was built to beat.

That matters for anyone building a business on a specific cost per token. Cerebras is the last large pure-play fast-inference vendor with its own silicon and its own incentive to keep prices down.

How much does fast AI inference cost per million tokens?

Cerebras is the most expensive way to run GPT-OSS-120B among mainstream providers. Groq sits mid-pack. Commodity GPU serverless tiers cost a fraction of both. The spread on the identical open-weights model is roughly 7x on output tokens, which is far wider than most teams assume.

Here is the pricing snapshot for GPT-OSS-120B as tracked by PricePerToken on August 21, 2026, with measured speed where it is published.

Provider Input $/M Output $/M 1M in + 1M out Measured output speed
Cerebras $0.35 $0.75 $1.10 1,641 tok/s
Groq $0.15 $0.60 $0.75 ~500 tok/s
SambaNova $0.14 $0.95 $1.09 Not published
Together AI $0.15 $0.60 $0.75 Not published
Amazon Bedrock $0.15 $0.60 $0.75 Not published
Baseten $0.10 $0.50 $0.60 Not published
Google $0.09 $0.36 $0.45 Not published
Fireworks $0.10 $0.10 $0.20 Not published
DeepInfra $0.037 $0.170 $0.207 Not published
OpenRouter $0.030 $0.170 $0.200 Not published
Cerebras speed from Artificial Analysis; Groq speed from CloudZero, May 2026.

Run a balanced million-in, million-out workload and Cerebras costs $1.10 against Groq’s $0.75 — a 47% premium. On output tokens alone the gap narrows to 25%.

Both are expensive next to Fireworks at $0.20 or the OpenRouter route at $0.20 for the same weights.

The discount lever most teams forget

Groq’s list price is not its real price. CloudZero notes that Groq’s Batch API and prompt caching each cut rates by 50%, and the two stack to roughly 25% of on-demand pricing.

Applied to GPT-OSS-120B, that pushes effective output cost toward $0.15 per million. Cerebras publishes no equivalent public stacking discount; its pricing page lists a $5 free trial, a $10 self-serve developer tier, and custom enterprise rates.

For any workload that tolerates a batch window, Groq is not 20% cheaper. It is closer to 5x cheaper.

Which is faster in practice, Cerebras or Groq?

Cerebras wins today, and not narrowly. Independent measurement puts Cerebras at 1,641 tokens per second on GPT-OSS-120B with 0.46 seconds to first token. Groq’s published figure for the same model is around 500 tokens per second. That is a 3.3x throughput advantage before the CS-4 ships at volume.

Artificial Analysis also clocks Cerebras at 1,402 tokens per second on Gemma 4 31B in reasoning mode, at a blended $0.24 per million.

What the speed gap costs in real money

Take a 100,000-token generation — a long agent trace or a full document rewrite.

  • Cerebras today: 61 seconds at 1,641 tok/s, costing $0.075 in output tokens.
  • Groq today: 200 seconds at 500 tok/s, costing $0.060.
  • Cerebras CS-4 claim: 23 seconds at 4,400 tok/s.
  • The math: 1.5 cents buys back 139 seconds — roughly 11 cents per minute of latency removed.

Eleven cents a minute is trivial if a customer is watching a cursor blink. It is indefensible if the job runs at 3 a.m. and nobody reads the output until morning.

One honest caveat: the 4,400 tok/s figure is a vendor claim tied to hardware that only started shipping this quarter. The 1,641 figure is measured on the live API. Do not budget against the former.

Is Cerebras worth the premium for AI agents?

Yes, for interactive and multi-step agent work — and this is the only case where the premium clearly pays. Agent loops multiply latency: ten sequential tool calls at 200 seconds each is a 33-minute task. The same loop at Cerebras speed finishes in about 10 minutes.

That compounding is the whole argument. A single completion at 500 tokens per second feels fine. Twenty of them chained behind a task does not.

Who should buy what

Use case Pick Why
Live chat, copilots, voice Cerebras 0.46s TTFT and 1,641 tok/s; latency is the product
Multi-step agents with tool calls Cerebras Per-step latency compounds across the loop
Overnight batch, evals, data labeling Groq (Batch API) Stacked discounts reach ~25% of list
High-volume, cost-capped production Fireworks / DeepInfra $0.20 per 1M in + 1M out on identical weights
Closed frontier models Neither Both serve open weights only
Sustained 24/7 single-model load Self-host Fixed GPU cost beats per-token above a break-even

That last row matters more as open weights close the quality gap. We ran the self-hosting break-even in our Qwen3.8-Max open weights breakdown, and the logic holds here.

Note the hard limit on both vendors: neither serves closed frontier models. If your stack depends on the newest proprietary coder, this comparison does not apply — see our GLM-5.3 vs DeepSeek V4 Pro comparison for the open-weight options that do run here.

What do the financials say about who wins?

Cerebras has the better technology and the shakier income statement. It went public on May 14, 2026, raising $5.5 billion. The stock priced at $185, opened at $385 for a 108% pop, and closed at $311 — a $66 billion valuation. Three months later the market repriced it hard.

The Q2 miss

On August 12, 2026, Cerebras reported Q2 revenue of $180.1 million against a $193.6 million consensus, per Investing.com. Adjusted EPS came in at a $2.98 loss versus an expected $0.18 loss. Shares fell 14% after hours.

Core revenue still grew 103% year over year to $209.9 million, and the company guided FY2026 core revenue to $880–890 million. Growth is not the problem. Margin is: guided core operating margin sits at negative 19% to negative 17%.

Why that should affect your buying decision

A vendor losing money on every wafer has two exits: raise prices or get acquired. Cerebras’ 2025 revenue was $510 million on 76% growth with $237.8 million of net income, so the balance sheet is not fragile — but the 2026 trajectory is being funded, not earned.

TrendForce values the company’s three-year OpenAI partnership at over $20 billion. That is concentration risk dressed as a moat.

The broader point TrendForce makes is the one to internalize: inference is a recurring cost tied directly to revenue, while training is a one-time R&D expense. Its example is brutal — Taalas’ HC1 delivers Llama 3.1 8B at 0.75 cents per million tokens against 3.79 cents on an Nvidia B200.

Specialized silicon is roughly five times cheaper per token than general-purpose GPUs at that scale. That is why the price you lock in today is unlikely to be the price in twelve months.

Frequently asked questions about Cerebras vs Groq

Is Cerebras faster than Groq?

Yes. Artificial Analysis measures Cerebras at 1,641 tokens per second on GPT-OSS-120B; Groq’s published figure for the same model is around 500. Cerebras also leads on time to first token at 0.46 seconds.

Is Groq cheaper than Cerebras?

Yes. On GPT-OSS-120B, Groq charges $0.15 input and $0.60 output per million tokens versus Cerebras at $0.35 and $0.75. With Groq’s stacked batch and cache discounts, the effective gap widens sharply.

Does Nvidia own Groq now?

Nvidia agreed in December 2025 to acquire Groq’s assets for about $20 billion, CNBC reported. The Groq API continues to operate and continues to publish its own pricing.

Can I run Claude or GPT-5 on Cerebras or Groq?

No. Both providers serve open-weights models only — GPT-OSS, Llama, Qwen, Gemma and similar. Closed frontier models stay on their vendors’ own APIs.

When does the CS-4 speed actually arrive for API users?

Cerebras says first CS-4 shipments began in Q3 2026. The 4,400 tokens per second per user figure is a vendor claim on new hardware, not yet an independently measured API result.

Is paying for faster inference ever worth it?

Only when latency is visible to a customer or compounds across an agent loop. At roughly 11 cents per minute of latency removed, speed is cheap for interactive products and pure waste for background jobs. We covered the same trade at the frontier-model layer in OpenAI Ultrafast vs Claude Fast Mode.

The bottom line

Route by whether something is waiting. If a human or an agent loop blocks on the token stream, Cerebras is worth its 25% output premium — 3.3x measured throughput for that price is one of the better deals in AI infrastructure, and the CS-4 should widen it.

If nothing is waiting, Cerebras is a rounding-error upgrade you are overpaying for. Send batch and evaluation traffic to Groq’s Batch API at roughly a quarter of list, or to Fireworks and DeepInfra at $0.20 per million in and out.

The strategic read is less comfortable. Nvidia owns Groq’s technology, Cerebras is losing money at negative 17% to 19% core operating margin, and specialized silicon is already showing five-fold cost advantages per token. Sign nothing longer than twelve months.

Sources

Comments

Leave a Reply

Discover more from Wealth Engine

Subscribe now to keep reading and get access to the full archive.

Continue reading