DeepSeek V4 Flash Beats Its Bigger Sibling for $0.14
2 min readChina’s DeepSeek has pushed its smallest frontier model into public beta, and the headline is not the architecture. It is that the little model now beats the much larger one in its own family on agent benchmarks, at a price that starts at 14 cents per million tokens.
What Shipped
DeepSeek released V4-Flash-0731 on July 31, 2026 and opened it on the API. The architecture is unchanged from the earlier V4-Flash preview: 284 billion total parameters with roughly 13 billion active, a one million token context window, and text only. The July 31 snapshot is a re-post-training update rather than a new design.
The gains landed where the industry is currently fighting hardest, which is agentic work. Terminal Bench 2.1 came in at 82.7, with Cybergym at 76.7, Toolathlon at 70.3 and DSBench-FullStack at 68.7. Those numbers now sit well above V4-Pro-Preview, the far larger model DeepSeek released earlier in the same family.
The Price Is the Point
Pricing held at $0.14 per million tokens for cache-miss input and $0.28 for output. Cached input runs around $0.003 per million, a discount of roughly 98 percent, which drags the blended rate down near $0.06 per million on workloads that reuse context heavily. Agent workloads reuse context constantly, so that cached rate is not a footnote.
For context on the gap: frontier proprietary models from the US labs are priced in dollars per million tokens, not cents. DeepSeek is also shipping this with open weights, so anyone can self-host it rather than pay per call at all.
Why It Matters
The pattern here is the one Chinese labs have been running all year. Rather than chase the top of the capability leaderboard, they target the price-performance curve underneath it and release the weights. Moonshot did it with Kimi K3 in July. DeepSeek is doing it now with a model small enough to run cheaply and good enough to drive tool use.
Watch whether Western providers respond on price, and whether V4 Flash holds up on independent agent evaluations rather than DeepSeek’s own. A small open model that genuinely outperforms its bigger sibling would say something uncomfortable about how much of frontier performance is really coming from scale.
