GLM-5.3-Flash is the cheapest 1M context model worth running in production. Z.ai lists it at $0.15 per million input tokens against $0.75 for Gemini 3.7 Flash — five times cheaper — while scoring 57 on the Artificial Analysis Intelligence Index versus Gemini’s 56. Google keeps two real advantages: raw throughput and vision. Everything else favors the open-weights challenger.
Z.ai shipped GLM-5.3-Flash on August 26, 2026, thirteen days after Google made Gemini 3.7 Flash generally available. Both models advertise a 1,048,576-token context window. Both target agentic coding and long-document work.
The gap is price. And at 1M-token scale, price is the entire product decision.
What is GLM-5.3-Flash?
GLM-5.3-Flash is a natively multimodal mixture-of-experts model with 320 billion total parameters and 18 billion active per token, released under an MIT license. It routes each token through 8 of 288 experts across 45 layers, ships in native FP8, and holds a 1,048,576-token context window.
That active-parameter count is the whole story. Z.ai is charging flagship-tier context for a model that only lights up 18B weights per forward pass.
The architecture behind the price
The model combines KDA linear-attention layers with NoPE sparse MLA layers. Per MarkTechPost’s launch coverage, that combination delivers roughly 3x less attention compute and a 4.4x smaller KV cache than GLM-5.3.
KV cache is what makes long context expensive to serve. Shrink it 4.4x and you can price a 1M window like a short one.
The jump over the previous generation is not cosmetic. Z.ai’s own numbers put DeepSWE v1.1 at 63.4%, up from 46.2% on GLM-5.2, and AutomationBench at 48.8%, up from 26.2% — a 22.6-point gain in one release cycle, according to LLM Stats.
How much does the cheapest 1M context model actually cost?
GLM-5.3-Flash lists at $0.15 per million input tokens and $0.50 output, with cached input at $0.03. Gemini 3.7 Flash lists at $0.75 input and $3.75 output on Google’s own model page. That is 5x on input and 7.5x on output, before any discount either side is running.
| Spec | GLM-5.3-Flash | Gemini 3.7 Flash |
|---|---|---|
| Released | Aug 26, 2026 | Aug 13, 2026 |
| License | MIT open weights | Proprietary API |
| Parameters | 320B total / 18B active | Undisclosed |
| Input context | 1,048,576 tokens | 1,048,576 tokens |
| Max output | 131,072 tokens | 65,536 tokens |
| Input / 1M | $0.15 | $0.75 |
| Output / 1M | $0.50 | $3.75 |
| Cached input / 1M | $0.03 | $0.06 (Vertex) |
| AA Intelligence Index | 57 | 56 |
| Output speed | 50.2 tok/s | 301 tok/s |
| Time to first token | 1.47s | 3.83s |
Pricing from Z.ai list rates and Google DeepMind’s Gemini Flash page. Speed and index figures from Artificial Analysis and Requesty’s Vertex listing. Resellers differ: OpenRouter lists GLM-5.3-Flash at $0.075 / $0.25 and Gemini 3.7 Flash at $0.375 / $1.875.
What it costs to fill the window once
Push a full 1,048,576-token context through each model, one time, and the arithmetic is brutal.
GLM-5.3-Flash: $0.157. Gemini 3.7 Flash: $0.786. Same window, same task, a $0.63 difference per call.
Run that 10,000 times a month — a modest document-processing pipeline — and you are looking at $1,573 versus $7,864. The $6,291 monthly delta is a headcount line item, not a rounding error.
For context on how wide the field has gotten, Morph’s context-window survey clocked a 71x spread between the cheapest and priciest 1M window on the market, from $0.14 on DeepSeek V4 Flash to $10.00 on Claude Fable 5.
The January 2027 price cliff
Google’s $0.75 / $3.75 is an introductory rate. Its own page states the promotion expires December 31, 2026, after which Gemini 3.7 Flash reverts to $1.50 per million input and $7.50 per million output.
On January 1, filling that same 1M window costs $1.57 on Gemini. Against GLM’s $0.157, that is a clean 10x.
Z.ai is running a promotion too — 50% off through September 9, 2026 — but its post-promo list price is the $0.15 already quoted. One vendor’s discount expires into a doubling. The other’s expires into the number on the page.
Which is better for coding agents, GLM-5.3-Flash or Gemini 3.7 Flash?
Gemini 3.7 Flash wins the coding benchmarks by margins too small to justify a 7.5x output bill. It leads Terminal-Bench 2.1 85.8% to 84.3% and DeepSWE v1.1 65.3% to 63.4%. GLM takes HLE 55.3% to 53.6% and destroys Gemini on AutomationBench, 48.8% to 30.4%.
A 1.5-point Terminal-Bench edge is inside the noise band of most agent harnesses. An 18.4-point AutomationBench gap is not.
AutomationBench measures multi-step tool use and workflow completion — the thing you actually buy an agent model for. GLM-5.3-Flash scores 60% higher there in relative terms.
Coding agents also burn output tokens, not input tokens. A long agentic run is thousands of generated tokens per step. That is precisely the axis where Gemini costs 7.5x more.
Our earlier breakdown of GLM-5.3 against DeepSeek V4 Pro found the same pattern in the open-weight tier: near-parity capability, order-of-magnitude price separation.
Where Gemini 3.7 Flash still wins
Google has genuine leads that no discount closes:
- Throughput: 301 tokens/second median output versus 50.2 for GLM-5.3-Flash — 6x faster generation.
- Vision: BabyVision 70.9% against GLM’s 53.4%, a 17.5-point gap that Z.ai does not dispute.
- Long-context recall: GDM-MRCR v2 at 128k scores 97.0%, among the strongest retrieval numbers published this year.
- Desktop agents: OSWorld-2.0 at 47.9% and Code Arena at 1588 Elo for web development.
- Vertical accuracy: Harvey LAB-AA at 90.7% on legal reasoning tasks.
GLM does answer faster on the first token — 1.47s versus 3.83s — which matters for interactive chat. But once generation starts, Gemini pulls away hard.
Is GLM-5.3-Flash worth it for multimodal work?
Only for charts and documents, not for general vision. GLM-5.3-Flash posts 78.0% on Chartography against DeepSeek-V4-Flash-Vision-Exp’s 64.3%, but trails Gemini 3.7 Flash badly on BabyVision, 53.4% to 70.9%. Structured visual data is a strength. Open-ended image understanding is not.
It also scores 62.4% on OfficeQA Pro, which points at the same conclusion: business documents, spreadsheets, slides and charts are where the multimodal stack earns its keep.
If your pipeline reads invoices, financial statements or dashboards, GLM handles it at a fifth of the price. If it captions arbitrary photos, pay Google.
We ran similar math on DeepSeek’s vision model against Claude Opus 4.8, where the per-image gap ran 23x. Cheap vision is now a solved category — you just have to match the model to the image type.
Which model should you buy for your workload?
Pick on token mix, not on leaderboard position. Input-heavy jobs at 1M scale go to GLM-5.3-Flash on cost alone. Latency-critical streaming and general vision go to Gemini 3.7 Flash. Coding agents are close on quality and lopsided on price.
| Use case | Buy | Why |
|---|---|---|
| Bulk document / RAG ingestion | GLM-5.3-Flash | $0.157 vs $0.786 per full 1M window |
| Long-horizon coding agents | GLM-5.3-Flash | AutomationBench 48.8 vs 30.4; 7.5x cheaper output |
| Real-time chat / streaming UX | Gemini 3.7 Flash | 301 tok/s vs 50.2 tok/s |
| General image understanding | Gemini 3.7 Flash | BabyVision 70.9 vs 53.4 |
| Charts, invoices, office docs | GLM-5.3-Flash | Chartography 78.0; OfficeQA Pro 62.4 |
| Needle-in-haystack retrieval | Gemini 3.7 Flash | GDM-MRCR v2 at 97.0% |
| Data that cannot leave your VPC | GLM-5.3-Flash | MIT weights, self-hostable |
| Terminal-Bench maximalists | Gemini 3.7 Flash | 85.8 vs 84.3 — for a 5x premium |
Should you self-host GLM-5.3-Flash instead?
Only above roughly 100 million tokens a month. The FP8 checkpoint is 306 GiB of weights before KV cache and needs NVIDIA Hopper or newer. That is a multi-GPU node running continuously against an API bill of $0.15 per million input tokens.
The MIT license is the real asset here, not the savings. It permits commercial use, modification and redistribution with no revenue thresholds — which is what makes GLM viable for regulated buyers who cannot route customer data through a third-party API.
Weights are published on Hugging Face as zai-org/GLM-5.3-Flash. Our Qwen3.8-Max self-hosting cost analysis laid out the crossover math in detail; the shape is unchanged, only the weight file got smaller.
Frequently asked questions
Is GLM-5.3-Flash actually the cheapest 1M context model?
Not quite. DeepSeek V4 Flash fills a 1M window for about $0.14 against GLM’s $0.157. But GLM scores 57 on the Artificial Analysis Intelligence Index and adds native multimodality, which makes it the cheapest capable one.
How much cheaper is GLM-5.3-Flash than Gemini 3.7 Flash?
Five times cheaper on input ($0.15 vs $0.75 per million) and 7.5 times cheaper on output ($0.50 vs $3.75). After Google’s introductory pricing expires December 31, 2026, the input gap widens to 10x.
Does GLM-5.3-Flash really have a 1M context window?
Z.ai specifies 1,048,576 tokens. OpenRouter lists its routed endpoint at 1,310,720 tokens with 131,072 max output — double Gemini 3.7 Flash’s 65,536-token output ceiling.
Which model is faster?
Gemini 3.7 Flash generates 6x faster at 301 tokens per second versus 50.2. GLM-5.3-Flash responds faster initially, at 1.47 seconds to first token against Gemini’s 3.83 seconds.
Are these benchmark scores independently verified?
Partly. The Artificial Analysis Intelligence Index scores are third-party. The Terminal-Bench, DeepSWE and AutomationBench figures are vendor self-reported on both sides — LLM Stats flags this explicitly for GLM-5.3-Flash.
Can I use GLM-5.3-Flash commercially?
Yes. The weights ship under an MIT license, which permits commercial use, modification and redistribution without revenue caps or usage restrictions.
What happens to Gemini 3.7 Flash pricing in 2027?
Google’s model page states the introductory rate ends December 31, 2026, moving to $1.50 per million input tokens and $7.50 per million output from January 1, 2027.
The bottom line
Buy GLM-5.3-Flash. For any workload dominated by input tokens or agent output tokens, it is the correct default — 5x to 7.5x cheaper at an intelligence index one point above Gemini 3.7 Flash, with a bigger output ceiling and weights you can take in-house.
Keep Gemini 3.7 Flash for exactly two jobs: user-facing streaming where 301 tokens per second is the product, and general image understanding where 17.5 BabyVision points decide whether the feature works at all.
The broader signal matters more than either model. Google discounted a Flash-tier model and still got undercut 5x by open weights released thirteen days later. Google’s own price sheet says that gap widens to 10x in four months.
If you are still routing 1M-token jobs through a proprietary Flash endpoint in 2027, you are paying a tenfold convenience tax. Compare that against our Gemini 3.7 Flash versus Claude Sonnet 5 cost-per-point analysis and the direction is unmistakable.
