Gemini 3.7 Flash vs Claude Sonnet 5 comes down to one number: cost per benchmark point. Google’s new workhorse matches Sonnet 5 on production coding evals while listing at $0.75/$3.75 per million tokens against Anthropic’s $2/$10. That is roughly 2.7x cheaper for equal-or-better coding output. Sonnet 5 still wins on the hardest reasoning tests. For agent workloads that burn tokens all day, Flash wins on money.
Google shipped Gemini 3.7 Flash on August 13, 2026 — three weeks after Gemini 3.6 Flash. The release matters less as a launch and more as a repricing event. When a cheap model closes the coding gap with a premium model, every AI budget line gets renegotiated.
We priced all three frontier options against their published benchmarks using vendor list prices. The result is not close.
How much does Gemini 3.7 Flash cost compared to Claude Sonnet 5?
Gemini 3.7 Flash lists at $0.75 per million input tokens and $3.75 per million output through December 31, 2026, per Google’s official Gemini API pricing page. Claude Sonnet 5 lists at $2.00 and $10.00. On an 80/20 input-output mix, that is $1.35 versus $3.60 per million tokens.
Anthropic also settled a question that had been hanging over Sonnet 5’s price. Its pricing documentation now states that the introductory $2/$10 rate is permanent and the scheduled September 1, 2026 increase to $3/$15 “will not occur.”
That was a defensive move. It did not close the gap.
The full price and spec comparison
| Model | Input / 1M | Output / 1M | Blended (80/20) | Context | Batch in / out |
|---|---|---|---|---|---|
| Gemini 3.7 Flash | $0.75 | $3.75 | $1.35 | 1M in / 64K out | $0.375 / $1.875 |
| Gemini 3.7 Flash (from Jan 1, 2027) | $1.50 | $7.50 | $2.70 | 1M in / 64K out | $0.75 / $3.75 |
| Claude Sonnet 5 | $2.00 | $10.00 | $3.60 | 1M at standard rate | $1.00 / $5.00 |
| GPT-5.6 Terra | $1.00 | $6.00 | $2.00 | Long-context tier priced separately | — |
| GPT-5.6 Sol | $2.50 | $15.00 | $5.00 | Long-context tier priced separately | — |
| Claude Opus 5 | $5.00 | $25.00 | $9.00 | 1M at standard rate | $2.50 / $12.50 |
The January 1, 2027 price cliff
Google’s $0.75 rate is introductory. On January 1, 2027 it doubles to $1.50/$7.50, which lifts the blended cost to $2.70.
Even then, Flash stays 25% under Sonnet 5. But the deepest discount window is four and a half months wide, and it is the single best arbitrage on the table right now.
The tokenizer tax nobody prices in
Anthropic’s documentation carries a note most comparison tables ignore: Claude 4.7 and later models use a newer tokenizer that “produces approximately 30% more tokens for the same text.”
Sticker price is per token. Your bill is per document. If that 30% applies to your workload, Sonnet 5’s effective cost per page of English moves closer to $4.70 blended — over 3x Flash’s introductory rate.
Gemini 3.7 Flash vs Claude Sonnet 5: which is better for coding?
Flash wins on production coding and agentic execution; Sonnet 5 wins on long-horizon reasoning. On Google’s published model-card comparisons, Flash takes FrontierCode 1.1 at 43.6% against Sonnet 5’s 42.7%, and crushes it on AutomationBench, 30.4% to 10.7%. Sonnet 5 answers on GDPval and Agent’s Last Exam.
The benchmark table below is vendor-stated from Google’s model card, tabulated independently by DataCamp and other outlets. Treat first-party numbers with the usual skepticism — but they are consistent across sources.
| Benchmark | Gemini 3.7 Flash | Claude Sonnet 5 | GPT-5.6 Terra |
|---|---|---|---|
| FrontierCode 1.1 (production code) | 43.6% | 42.7% | 41.3% |
| DeepSWE v1.1 (long-horizon SWE) | 65.3% | 53.8% | 69.6% |
| Terminal-bench 2.1 | 85.8% | 80.4% | 87.4% |
| WebDev Arena (Elo) | 1588 | 1541 | 1523 |
| AutomationBench | 30.4% | 10.7% | 23.6% |
| GDPval-AA v2 (Elo) | 1525 | 1598 | 1578 |
| GDM-MRCR v2, 128k (recall) | 97.0% | 81.5% | 93.5% |
| Agent’s Last Exam | 26.3% | 33.3% | 28.0% |
Cost per benchmark point: the number that decides it
Divide blended cost by score and the argument ends.
- FrontierCode 1.1: Flash costs $0.031 per point. Sonnet 5 costs $0.084. Terra costs $0.048.
- DeepSWE v1.1: Flash $0.021 per point, Terra $0.029, Sonnet 5 $0.067.
- AutomationBench: Flash delivers 2.8x Sonnet 5’s score at 37% of the price.
- Long-context recall (MRCR 128k): Flash leads by 15.5 points and costs 63% less.
Sonnet 5 is charging a 2.7x premium to lose a coding benchmark by 0.9 points. That is not a defensible position in a procurement meeting.
Where Claude Sonnet 5 still earns its price
Two places. Sonnet 5 leads GDPval-AA v2 at 1598 Elo against Flash’s 1525 — that benchmark tracks economically valuable knowledge work, not code. It also leads Agent’s Last Exam, 33.3% to 26.3%.
If your workload is legal analysis, financial modeling, or research synthesis rather than shipping code, the premium is arguable. If it is code, it is not.
How does GPT-5.6 Terra change the math?
Terra is the quiet value play, and most launch-day comparison tables mispriced it. OpenAI’s official pricing page lists gpt-5.6-terra at $1.00 input and $6.00 output per million, with a separate long-context tier at $2.00/$9.00 — not the $2.00/$12.00 figure that circulated all week.
At the correct list price, Terra blends to $2.00 per million. It also posts the best DeepSWE v1.1 score in the group at 69.6% and the best Terminal-bench 2.1 at 87.4%.
Terra’s cached input runs $0.10 per million, half of Anthropic’s $0.20 cache-hit rate for Sonnet 5. For retrieval-heavy agents replaying the same system prompt thousands of times a day, that difference compounds fast.
What does this actually cost at production volume?
Take a mid-size agent workload: 500 million input tokens and 100 million output tokens per month. That is a realistic footprint for a coding assistant serving a 50-engineer team. The spread between the cheapest and most expensive option is $4,250 a month.
| Model | Monthly cost | Annualized | vs. Sonnet 5 |
|---|---|---|---|
| Gemini 3.7 Flash (intro) | $750 | $9,000 | −$15,000/yr |
| GPT-5.6 Terra | $1,100 | $13,200 | −$10,800/yr |
| Gemini 3.7 Flash (2027 rate) | $1,500 | $18,000 | −$6,000/yr |
| Claude Sonnet 5 | $2,000 | $24,000 | — |
| GPT-5.6 Sol | $2,750 | $33,000 | +$9,000/yr |
| Claude Opus 5 | $5,000 | $60,000 | +$36,000/yr |
Batch processing cuts all of it roughly in half. Gemini 3.7 Flash drops to $0.375/$1.875 through year-end; Sonnet 5 drops to $1.00/$5.00. The ranking does not change.
Which model should you use for which job?
Match the model to the failure mode you can least afford. Coding agents that run unsupervised for hours need throughput and cheap retries. Client-facing analysis needs reasoning depth. Nothing here is a universal answer, and paying Opus prices for autocomplete is how AI budgets die.
| Use case | Best choice | Why |
|---|---|---|
| High-volume coding agents | Gemini 3.7 Flash | Top FrontierCode score at 37% of Sonnet 5’s blended price |
| Long-horizon autonomous SWE | GPT-5.6 Terra | Leads DeepSWE (69.6%) and Terminal-bench 2.1 (87.4%) |
| Front-end and web generation | Gemini 3.7 Flash | WebDev Arena Elo 1588, ahead of both rivals |
| Legal, financial, research synthesis | Claude Sonnet 5 | Top GDPval-AA v2 Elo at 1598 |
| Million-token document pipelines | Gemini 3.7 Flash | 97.0% MRCR recall at 128k, cheapest per token |
| Hardest reasoning, cost no object | Claude Opus 5 | Frontier tier — $9.00 blended, use sparingly |
| Bulk offline processing | Gemini 3.7 Flash (batch) | $0.375 / $1.875 through Dec 31, 2026 |
Is switching to Gemini 3.7 Flash worth it in 2026?
Yes, if your token spend clears roughly $1,000 a month. Below that, migration engineering costs more than it saves. Above it, the savings compound — and Google’s three-week release cadence means the model you migrate to keeps improving without a renegotiation.
The strategic read is bigger than one model. Frontier-tier coding capability is commoditizing on a quarterly clock, and price is the only lever customers can still feel. We saw the other side of that trade this week when DeepSeek raised prices by up to 1,100% overnight — the cheap-inference era is being rationed, not extended.
Anthropic is spending to stay in the fight. It reached a $2 trillion valuation and put $6 billion into getting cheaper. Cancelling the September price increase is the visible half of that strategy.
One caution before you point an agent at production: capability and autonomy scale together. The same agentic execution that makes Flash cheap to run is what let an AI agent crack 85 accounts in four days. Sandbox accordingly.
Frequently asked questions
Is Gemini 3.7 Flash actually cheaper than Claude Sonnet 5?
Yes. $0.75/$3.75 per million tokens versus $2.00/$10.00 — about 2.7x cheaper on an 80/20 blend. The introductory rate holds through December 31, 2026, then doubles to $1.50/$7.50.
Does Gemini 3.7 Flash beat Claude Sonnet 5 at coding?
On Google’s published card, yes — narrowly on FrontierCode 1.1 (43.6% vs 42.7%), decisively on DeepSWE v1.1 (65.3% vs 53.8%) and AutomationBench (30.4% vs 10.7%). GPT-5.6 Terra still leads DeepSWE overall at 69.6%.
What is Gemini 3.7 Flash’s context window?
One million input tokens and up to 65,536 output tokens, with a March 2026 knowledge cutoff. Claude Sonnet 5 also offers a 1M-token window at standard per-token pricing.
Did Claude Sonnet 5’s price go up on September 1, 2026?
No. Anthropic’s documentation confirms the scheduled increase to $3/$15 per million tokens will not occur. The $2/$10 introductory rate is now the standard price.
How much does GPT-5.6 Terra cost?
$1.00 input and $6.00 output per million tokens on the standard tier, with a long-context tier at $2.00/$9.00. Cached input is $0.10 per million.
Where can I use Gemini 3.7 Flash today?
Google AI Studio, Android Studio, Google Antigravity, the Gemini Enterprise Agent Platform, and Gemini Spark for AI Pro and Ultra subscribers across 160+ countries.
Should I run one model or mix them?
Mix. Route bulk coding and document work to Flash, long-horizon autonomous tasks to Terra, and high-stakes analysis to Sonnet 5. Routing the majority of calls to the cheapest capable tier is where the savings actually come from.
The bottom line
Default to Gemini 3.7 Flash for coding and agent workloads. It wins or ties on the coding benchmarks that map to shipped software, costs $1.35 blended against Sonnet 5’s $3.60, and saves a 50-engineer team roughly $15,000 a year at the volumes above.
Keep Claude Sonnet 5 for the narrow band where it leads: GDPval-style knowledge work and Agent’s Last Exam reasoning. Keep GPT-5.6 Terra for long-horizon autonomous engineering, where its 69.6% DeepSWE score is worth the extra $0.65 per million blended.
And put a calendar reminder on December 31, 2026. That is when Google’s discount ends and this entire calculation gets re-run. The labs are shipping every three weeks now — the money chasing this market guarantees the next repricing is already in the pipeline.
Leave a Reply