Eleven hundred percent.
That is the top-end figure buried in the DeepSeek price increase that takes effect on Sunday, August 16, 2026 — and it comes from the one company in artificial intelligence whose entire global reputation was built on being impossibly, almost suspiciously cheap.
For eighteen months, DeepSeek was the argument. Every time someone said frontier AI was structurally expensive, someone else pointed at Hangzhou and said: no, it isn’t — they’re doing it for pennies. That argument moved markets. It rewrote capex assumptions. It made a generation of investors believe inference costs would fall forever, like transistors, like bandwidth, like everything else in tech.
On August 13, DeepSeek shipped its flagship DeepSeek V4-Pro to general availability. Three days later, it is quadrupling the price of running it.
The direction of travel just reversed. And the reason it reversed is the most important thing in this story.
What the DeepSeek price increase actually changes
Strip out the percentages and look at the raw per-token numbers, because the percentages are doing a lot of theatrical work.
For V4-Pro, output tokens go from a flat $0.87 per million to $3.96 per million during peak hours — roughly a 4.5x jump — and $1.98 per million off-peak. Cache-miss input tokens move from $0.435 per million to $1.32 peak and $0.66 off-peak.
For the cheaper V4-Flash tier, output goes from $0.28 per million to $1.32 peak and $0.66 off-peak. Cache-miss input rises from $0.14 to $0.44 peak and $0.22 off-peak.
The headline 1,100% figure comes from the cached input tier — the deeply discounted rate DeepSeek charged when a prompt prefix was already sitting in its KV cache. That was the single cheapest number in commercial AI, and it is where the proportional increase is most violent. Reported increases across the cached tier run from roughly 52% to 1,100%, depending on model and time of day.
The peak/off-peak fine print nobody put in the headline
DeepSeek did not simply raise a number. It introduced time-of-day pricing, which is a structurally different product.
Peak windows are 01:00–04:00 and 06:00–10:00 UTC. Everything outside those seven hours is off-peak, billed at exactly half the peak rate. The company framed the change in its developer documentation as an effort to allocate resources “more reasonably” and to nudge batch workloads into quieter hours.
That framing matters. As one analyst quoted by InfoWorld put it, 17 of every 24 hours stay at half price. A team running overnight evaluation sweeps, document ingestion, or scheduled agent runs can absorb most of this with a cron change. A team serving live user traffic in Asian business hours cannot.
Utilities price electricity by time of day because generation capacity is finite. DeepSeek just did the same thing to tokens.
Why DeepSeek raising prices matters more than the percentage
There is a detail here that is easy to skim past and shouldn’t be.
DeepSeek’s rock-bottom rates were originally promotional, scheduled to expire on May 31. The company then announced it was making those discounted rates permanent. It has now reversed that decision inside a single quarter.
Companies do not walk back a public permanence commitment on pricing because things are going well. They do it because the unit economics moved underneath them. DeepSeek’s own stated reason — resource allocation — is a polite way of saying demand is outrunning the compute it can get its hands on.
Reporting on the change from InfoWorld and Computerworld framed it exactly that way: prices are rising because AI demand is straining capacity. The analyst quote is almost aggressively simple: “when demand goes up, pricing goes up, because supply becomes constrained.”
That is the part with implications far beyond one Chinese lab. The entire bull case for cheap AI has rested on an assumption that inference is a software problem that gets cheaper on a curve. What August 16 suggests is that inference is a power and silicon problem, and those curves behave differently. It is the same pressure driving Anthropic to spend $6 billion buying its way to cheaper inference rather than waiting for hardware to save it, and the same pressure behind five companies committing $650 billion of capital expenditure in a single year.
DeepSeek has an additional constraint its Western competitors do not share: export controls. It cannot simply write a larger check to Nvidia. When a lab that cannot buy its way out of a capacity crunch starts rationing by price, that is a supply signal, not a greed signal.
What DeepSeek V4-Pro is — and what nobody has independently verified
The model itself is not an afterthought. V4-Pro is reportedly a 1.6-trillion-parameter mixture-of-experts system that activates only about 49 billion parameters per token — which is precisely how the old $0.87 output price was possible at all.
DeepSeek’s own reported gains over its April preview build are large:
- DeepSWE (software engineering): 12.8 → 62.7
- CyberGym (vulnerability discovery): 52.7 → 83.3
- DSBench-Hard (data science): 31.1 → 67.2
- Terminal Bench 2.1 (agentic terminal use): 87.9
- Humanity’s Last Exam: 42.7 out of a reported 60.0 ceiling
The release also adds three “thinking effort” levels — low, high and max — and native Responses API support so V4-Pro can be dropped into Codex-style tooling. Alongside it, DeepSeek shipped a developer preview of DeepSeek Harness, an agentic coding harness positioned as an open competitor to Claude Code.
Now the caveat, and it is a real one: as of publication, no third-party evaluator has replicated those scores. DeepSeek has not published the evaluation harness used to produce them. Treat every number above as a vendor claim until someone independent runs it.
The CyberGym figure deserves particular scrutiny given how quickly frontier models are being pointed at security work — a trajectory we covered when OpenAI’s security model surfaced live Chrome vulnerabilities. A self-reported 83.3 on vulnerability discovery is either a significant capability milestone or a benchmark artifact, and right now there is no way to tell which.
There is also a governance dimension for regulated buyers. DeepSeek’s hosted API operates under Chinese law, and no named independent security audit of V4-Pro’s weights has been published. For a US bank or hospital system, that is a procurement blocker regardless of price.
Google went the opposite direction on exactly the same day
Here is the contradiction that makes this week genuinely strange.
On August 13 — the same day DeepSeek’s V4-Pro went GA with a price hike queued behind it — Google launched Gemini 3.7 Flash and cut the price in half.
Introductory pricing is $0.75 per million input tokens and $3.75 per million output tokens, running through December 31, 2026. The model keeps a roughly 1,048,576-token context window with a 65,536-token output limit, and posts substantial coding gains: DeepSWE v1.1 from 49.0% to 65.3%, FrontierCode 1.1 from 34.4% to 43.6%, AutomationBench from 17.0% to 30.4%, and a 1,588 rating on WebDev Arena.
Read the fine print, though. That discount has an expiry date. On January 1, 2027, the list price reverts to $1.50 and $7.50 — double. Google isn’t claiming a permanent cost breakthrough. It is running a limited-time land grab and telling you so in the terms.
OpenAI did something structurally similar in late July, cutting GPT-5.6 Luna’s price by roughly 80% as enterprise buyers grew visibly cost-sensitive.
So the picture is not “AI is getting more expensive.” The picture is: the players with hyperscale balance sheets and their own data centers are still buying market share with subsidized tokens, and the player without those things just stopped being able to.
The money: who actually eats a 4x inference bill
Percentage increases land unevenly, and the distribution is the story.
The hardest hit are the businesses whose entire margin structure was underwritten by DeepSeek’s cached-input rate: retrieval-heavy products that stuff the same 100,000-token corpus into every request, AI wrapper startups whose pricing pages promise unlimited usage, and agentic products that burn output tokens in long reasoning chains. A 4.5x output increase against a fixed subscription price is not a cost problem; it is a business model problem.
Least affected are batch-tolerant enterprises. Overnight ETL, nightly code review, offline document classification — all of it can be scheduled into the 17 off-peak hours, where the effective increase is roughly half the headline.
Quietly advantaged: Google, OpenAI and Anthropic. Every enterprise procurement team that built a cost model on DeepSeek’s permanence promise now has to rebuild it, and rebuilding is when vendors get switched. The context here is worth remembering — this is the same market where developers are already paying $200 a month for frontier access and questioning what they get for it.
Even after the increase, DeepSeek is not expensive in absolute terms. Comparable output pricing at Moonshot’s Kimi K3 has been reported around $15 per million tokens and OpenAI’s GPT-5.6 Sol around $30, with premium Anthropic tiers reported far higher still. DeepSeek’s $3.96 peak remains an order of magnitude below the top of the market. But OpenAI’s budget GPT-5.6 Luna reportedly undercuts DeepSeek’s Flash tier at peak — and that is new. For the first time, the cheap-tier crown is contested.
The counterargument: this may be less apocalyptic than it looks
Honesty requires acknowledging that “1,100%” is the most misleading number in this story.
It applies to the cached-input tier, the smallest line item on most bills, and only at peak. On blended real-world workloads, most teams will see something closer to a 2x to 3x increase — meaningful, but not existential, and starting from a base so low that the absolute dollars are still small for anyone below serious scale.
Second, off-peak pricing is a genuine option, not a rhetorical dodge. Seventeen hours a day at half price is a real lever for anyone whose latency requirements are loose.
Third, and most importantly: a company raising prices during a capacity crunch is behaving rationally, not desperately. Underpricing scarce compute produces queueing, degraded latency and outages. Price is the least bad rationing mechanism available. There is a plausible reading in which this is a sign of demand strength, not weakness.
And a fourth caveat worth stating plainly: DeepSeek has not published audited unit economics. Nobody outside the company knows whether the old prices were near cost, deeply subsidized, or somewhere in between. Anyone telling you they know what this proves about the true cost of inference is guessing.
What to watch next
- Independent V4-Pro benchmarks. If outside evaluators reproduce the DeepSWE and CyberGym numbers, the price increase looks like confident pricing of a genuinely strong model. If they don’t, it looks like margin defense wrapped in a launch.
- Whether rivals follow. If Alibaba’s Qwen, Moonshot or Z.ai raise prices in the next 60 days, the Chinese AI price war is structurally over. If they hold and take share, DeepSeek’s move looks idiosyncratic.
- January 1, 2027. The date Gemini 3.7 Flash reverts to $1.50 / $7.50. If Google extends the discount, the subsidy war continues. If it lets the price double, the cheap-inference era has an official end date.
- Off-peak utilization data. If DeepSeek’s peak windows stay saturated even after the price change, the capacity constraint is worse than disclosed.
- Enterprise churn. Watch whether OpenRouter and similar aggregators report traffic shifting away from DeepSeek endpoints after August 16.
Bottom line
The DeepSeek price increase is not the story because of the number. It is the story because of the direction.
For two years the industry has operated on an unexamined assumption that the cost of intelligence falls monotonically. This week, the company that did the most to popularize that assumption broke its own permanence pledge and started charging by the hour — the way you charge for electricity, not the way you charge for software.
Google’s simultaneous half-price launch doesn’t refute that. It reinforces it. When only companies with their own data centers can afford to keep cutting, cheap AI stops being a technology trend and becomes a balance-sheet privilege.
Frequently Asked Questions
How much is the DeepSeek price increase?
It varies by tier. V4-Pro output rises from $0.87 to $3.96 per million tokens at peak and $1.98 off-peak. V4-Flash output rises from $0.28 to $1.32 peak and $0.66 off-peak. Cache-miss input roughly doubles to triples. The widely quoted 1,100% figure applies to the cached-input tier at peak hours, which is the smallest component of most bills — blended real-world increases are typically closer to 2x–3x.
When does the new DeepSeek API pricing take effect?
The new rates take effect on Sunday, August 16, 2026, at 16:00 UTC, according to DeepSeek’s developer documentation. The change applies to both V4-Pro and V4-Flash on the hosted API. Existing integrations do not need code changes; the same model endpoints simply bill at the new peak and off-peak rates from that timestamp forward.
What are DeepSeek’s peak and off-peak hours?
Peak windows are 01:00–04:00 UTC and 06:00–10:00 UTC — seven hours total. Every other hour of the day is off-peak and billed at exactly half the peak rate. That leaves 17 of 24 hours at the discounted rate, which is why batch-tolerant workloads such as overnight evaluations, document ingestion and scheduled agent runs can absorb much of the increase by rescheduling.
Is DeepSeek still cheaper than OpenAI and Anthropic?
At the frontier tier, yes, and by a wide margin. DeepSeek V4-Pro’s $3.96 peak output price sits far below reported list rates for OpenAI’s GPT-5.6 Sol and premium Anthropic tiers. The exception is the budget segment: OpenAI’s GPT-5.6 Luna, cut roughly 80% in late July, reportedly undercuts DeepSeek’s V4-Flash at peak hours. That is the first serious challenge to DeepSeek’s cheap-tier position.
What is DeepSeek V4-Pro?
DeepSeek V4-Pro is the company’s flagship model, released to general availability on August 13, 2026. It is reportedly a 1.6-trillion-parameter mixture-of-experts architecture activating roughly 49 billion parameters per token, with three thinking-effort levels and native Responses API support. DeepSeek reports large agentic and coding gains, but no independent evaluator has replicated those benchmark scores as of publication.
Why is DeepSeek raising prices?
DeepSeek says the goal is to allocate resources more reasonably by shifting flexible workloads into off-peak hours. Industry reporting frames it as a capacity constraint: demand for agentic and reasoning workloads is growing faster than available compute, and export controls limit how quickly DeepSeek can add hardware. Time-of-day pricing is a rationing mechanism, the same tool utilities use for electricity.
Sources
- Caixin Global — DeepSeek Launches V4-Pro and Raises API Prices by as Much as 1,100%
- DeepSeek API Docs — Change Log (V4-Pro GA and pricing adjustment)
- Engadget — DeepSeek’s AI models are about to cost four times more
- Quartz — DeepSeek raising API prices by up to 1,100% starting Aug. 16
- InfoWorld — DeepSeek raises some V4 prices by more than 10x as AI demand strains capacity
- Fortune — DeepSeek increases prices for AI services by multiple times
- Tech Times — DeepSeek V4-Pro 0813 goes GA; benchmark claims await independent proof
- VentureBeat — Google’s Gemini 3.7 Flash targets coding and agents with a 50% introductory price cut
- Android Headlines — Gemini 3.7 Flash launches with major performance gains and half-off pricing
- CNBC — OpenAI cuts prices for two of its GPT-5.6 AI models
Disclaimer: This article is journalism, not investment advice. It discusses company pricing, valuations and market dynamics for informational purposes only. Figures are as reported at the time of publication and may change. Nothing here is a recommendation to buy, sell or hold any security. Do your own research and consult a licensed financial professional before making investment decisions.

Leave a Reply