DeepSeek V4, explained
Facts checked 2026-07-21DeepSeek V4 arrived on 2026-04-24 as two open-weight models under MIT license: V4 Pro (1.6 trillion parameters, 49B active — the largest open-weight model ever released) and V4 Flash (284B/13B), both with a million-token context. Three months later, Flash is the single most-used model on OpenRouter at 5.3 trillion tokens a week. This is the lab that crashed Nvidia’s stock once, and what it built since.
The quant fund that broke the market
DeepSeek grew out of High-Flyer, a Hangzhou quantitative hedge fund co-founded by Liang Wenfeng, which had been stockpiling GPU clusters for trading research years before spinning up an AGI lab in 2023. The team stayed famously small — around 160 people — and hired for raw ability over experience (Wikipedia).
Two moments made it a household name. First, the V3 technical report priced its final training run at $5.576M — a number that reframed the economics of frontier AI even after critics noted it excluded the ~$1B+ of infrastructure behind it. Then R1 (January 2025) hit #1 on the US App Store and wiped over $500B off Nvidia in a single day — the largest one-day value loss in market history. By July 2026 the arc had fully inverted: Reuters reports a raise at a $74B valuation ahead of an IPO, on annualized revenue approaching $500M.
From R1 to V4
The famous "R2" never shipped — reporting attributes the delay to failed training runs on domestic Huawei silicon before a return to Nvidia hardware. Instead, DeepSeek folded reasoning into a hybrid line (V3.1, then V3.2 with its sparse-attention experiment and a 50%+ price cut), and V4 completed the merge: one model family, three reasoning modes (non-think / think-high / think-max), OpenAI and Anthropic API formats, and 1M-token context as the default everywhere. The legacy deepseek-chat and deepseek-reasoner aliases retire on 2026-07-24.
Inside the models
- V4 Pro: 1.6T total / 49B active MoE. V4 Flash: 284B / 13B — reasoning that “closely approaches” Pro, per DeepSeek, at a fraction of the cost.
- Compressed Sparse Attention: at 1M context, V4 Pro needs only 27% of the inference FLOPs and 10% of the KV cache of its predecessor — the technique lineage that started with V3.2’s DSA.
- Pre-trained on 32T+ tokens; post-trained by cultivating domain-expert models with RL, then distilling them back into one network. Weights ship in FP4+FP8.
- Both models are text-only — no vision. That, and knowledge-recall gaps flagged by independent evals, are the honest limitations.
What the benchmarks say
| Benchmark | V4 Pro | V4 Flash | Context |
|---|---|---|---|
| LiveCodeBench | 93.5 | 91.6 | Best in DeepSeek’s comparison table, above Gemini-3.1-Pro |
| Codeforces rating | 3206 | 3052 | Above GPT-5.4 (3168) |
| SWE-bench Verified | 80.6 | 79.0 | Claude Opus 4.6: 80.8 — effectively tied |
| GPQA Diamond | 90.1 | 88.1 | Gemini-3.1-Pro: 94.3 |
| MMLU-Pro | 87.5 | 86.2 | Ties GPT-5.4 |
| AA Intelligence Index (independent) | 52 | 47 | #2 open-weights family behind Kimi K2.6; Flash ≈ Sonnet-4.6 class |
| SWE-bench Verified (NIST CAISI, independent) | 74% | — | GPT-5.5: 81% — CAISI places V4 ~8 months behind the US frontier |
Self-reported numbers from the V4 Pro model card; independent rows from Artificial Analysis and the NIST CAISI evaluation. AA flags heavy verbosity and weak factual recall (its Omniscience test) — worth knowing for knowledge-heavy workloads.
The pricing story
V4 launched with Pro at a list price of $1.74/$3.48 per million tokens, immediately “discounted 75%” to $0.435/$0.87 — a discount that quietly became the standard price by June. Flash sits at $0.14/$0.28, undercutting even DeepSeek’s previous V3.2 while multiplying context 8×. Cache hits are absurdly cheap ($0.0028–$0.0036 per million — up to 1/120th of a miss). Next up, per TechNode: the full V4 release introduces DeepSeek’s first peak/off-peak pricing, with peak hours billed at double the off-peak rate. Smouter’s live rate for deepseek-v4-pro right now: $0.217 in / $0.435 out — 50% under DeepSeek’s standard price, all hours, no peak windows.
Adoption, and the fine print
By July 2026, V4 Flash alone moved more than double the weekly tokens of the most-used frontier model on OpenRouter — while costing ~23× less per token on average. The fine print from three months of production use: recurring peak-time 429/503 “server busy” responses on the first-party API, dynamic rate limits you cannot buy your way past, and hands-on tests (like Kilo’s build-a-project rubric) that found it capable but imperfect — Pro scored between Claude Opus and Kimi, and Flash built a near-working project for two cents.
Sources
- DeepSeek — official V4 preview announcement (2026-04-24)
- Hugging Face — DeepSeek-V4-Pro model card (MIT, specs, benchmarks)
- DeepSeek — live API pricing
- arXiv — DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence
- Artificial Analysis — “DeepSeek is back among the leading open-weights models”
- NIST CAISI — independent evaluation of DeepSeek V4 Pro (2026-05)
- TechCrunch — open-source volume vs frontier spend (OpenRouter/Vercel data)
- TechNode — full V4 launch and peak/off-peak pricing plan
- Kilo — hands-on V4 Pro & Flash build test
- Wikipedia — DeepSeek (company history, R1 market moment, V3 cost)