August 15, 2026

AIincider

AI News. No Noise. Just Signal.

DeepSeek Ends the Price War: API Costs Jump Up to 1,100%

2 min read
DeepSeek is raising V4 API prices by 50% to over 1,100% from August 16, adding peak and off-peak billing as demand strains capacity. Read the breakdown.

The company that started the AI price war just called it off. DeepSeek is raising what developers pay for its V4 models by anywhere from 50 percent to more than 1,100 percent, with the new rates going live at 16:00 UTC on August 16. The Chinese lab is also abandoning flat pricing entirely in favor of a peak and off peak structure, according to Fortune.

How DeepSeek Got Here

DeepSeek built its reputation on undercutting everyone. When R1 arrived in January 2025 at a fraction of the cost of comparable Western models, it reset expectations across the industry and forced OpenAI, Google and Anthropic to defend their pricing. The V4 family continued that strategy, and V4-Flash in particular became a default choice for developers running high volume workloads on thin margins.

Rock bottom prices only work if capacity is cheap. China’s leading labs are buying compute in a constrained market, with export controls limiting access to the highest end accelerators and domestic alternatives still ramping.

What the DeepSeek Price Increase Looks Like

V4-Flash output tokens move to $1.32 per million during peak hours and $0.66 per million off peak, up from a flat $0.28 per million. That is a 371 percent jump at peak. Cache miss input tokens rise to $0.44 per million at peak from $0.14, roughly a 214 percent increase. V4-Pro sees the steepest changes, reaching past 1,100 percent on some token types.

Peak hours are defined as 01:00 to 04:00 and 06:00 to 10:00 UTC, with off peak rates set at half the peak level. DeepSeek framed the change around demand straining capacity, which is an unusually blunt admission for a company that has traded on affordability.

Why It Matters

Cheap Chinese inference has been one of the strongest deflationary forces in AI, and its removal changes the math for every startup that built on it. Developers who chose V4-Flash purely on cost now have to re-run the comparison against Gemini Flash, GPT-5.6 Luna and the open weight Qwen models they can host themselves.

The broader signal is that inference capacity, not model quality, is becoming the scarce good. Watch whether Alibaba and Moonshot follow with their own increases. If they do, the era of racing prices toward zero is over.

Continue Reading…

Leave a Reply