Skip to content
DeepSeek V4 ProDeepSeek listSmouter (cached)You save
Input / 1M tokens$0.435$0.217−50%
Output / 1M tokens$0.87$0.435−50%
DeepSeek V4 FlashDeepSeek listSmouter (cached)You save
Input / 1M tokens$0.14$0.047−67%
Output / 1M tokens$0.28$0.094−67%

Smouter price is the live value-lane rate (cached) for deepseek-v4-pro and family. Reference is DeepSeek's published API list price (V4 Pro $0.435 in / $0.87 out, V4 Flash $0.14 in / $0.28 out per 1M tokens, cache-miss, July 2026). Prices shown per 1M tokens.

One line to switch

Smouter speaks the OpenAI API. If your app already calls DeepSeek V4 Pro through the vendor, OpenRouter, or the OpenAI SDK, you only change the base URL and the key — the request body is identical.

  • Drop-in chat/completions — streaming, tools, JSON mode
  • Same weights DeepSeek serves — Pro for hard problems, Flash for volume
  • Works in any coding agent or IDE that takes an OpenAI base URL
deepseek_v4.py
from openai import OpenAI client = OpenAI(    base_url="https://api.smouter.ai/v1",  # the only line you change    api_key="sk-smouter-...",) resp = client.chat.completions.create(    model="deepseek-v4-pro",  # same model, about half the price    messages=[{"role": "user", "content": "Explain this stack trace."}],)

Both V4s, one key

V4 Pro when the problem is hard, V4 Flash when the volume is high — switch between them with a string. The same key also serves R1, V3.2 and 100+ other models.

Routed below list, live

Smouter routes every request to the cheapest healthy provider serving the model and reprices continuously — the table above is the live rate, not a promo.

Pay per token, no lock-in

No subscription, no minimums, no card to start. Add credits, get an API key, and you are routing DeepSeek V4 in minutes. Cancel any time.

DeepSeek V4, explained

Facts checked 2026-07-21

DeepSeek V4 arrived on 2026-04-24 as two open-weight models under MIT license: V4 Pro (1.6 trillion parameters, 49B active — the largest open-weight model ever released) and V4 Flash (284B/13B), both with a million-token context. Three months later, Flash is the single most-used model on OpenRouter at 5.3 trillion tokens a week. This is the lab that crashed Nvidia’s stock once, and what it built since.

The quant fund that broke the market

DeepSeek grew out of High-Flyer, a Hangzhou quantitative hedge fund co-founded by Liang Wenfeng, which had been stockpiling GPU clusters for trading research years before spinning up an AGI lab in 2023. The team stayed famously small — around 160 people — and hired for raw ability over experience (Wikipedia).

Two moments made it a household name. First, the V3 technical report priced its final training run at $5.576M — a number that reframed the economics of frontier AI even after critics noted it excluded the ~$1B+ of infrastructure behind it. Then R1 (January 2025) hit #1 on the US App Store and wiped over $500B off Nvidia in a single day — the largest one-day value loss in market history. By July 2026 the arc had fully inverted: Reuters reports a raise at a $74B valuation ahead of an IPO, on annualized revenue approaching $500M.

From R1 to V4

The famous "R2" never shipped — reporting attributes the delay to failed training runs on domestic Huawei silicon before a return to Nvidia hardware. Instead, DeepSeek folded reasoning into a hybrid line (V3.1, then V3.2 with its sparse-attention experiment and a 50%+ price cut), and V4 completed the merge: one model family, three reasoning modes (non-think / think-high / think-max), OpenAI and Anthropic API formats, and 1M-token context as the default everywhere. The legacy deepseek-chat and deepseek-reasoner aliases retire on 2026-07-24.

Inside the models

  • V4 Pro: 1.6T total / 49B active MoE. V4 Flash: 284B / 13B — reasoning that “closely approaches” Pro, per DeepSeek, at a fraction of the cost.
  • Compressed Sparse Attention: at 1M context, V4 Pro needs only 27% of the inference FLOPs and 10% of the KV cache of its predecessor — the technique lineage that started with V3.2’s DSA.
  • Pre-trained on 32T+ tokens; post-trained by cultivating domain-expert models with RL, then distilling them back into one network. Weights ship in FP4+FP8.
  • Both models are text-only — no vision. That, and knowledge-recall gaps flagged by independent evals, are the honest limitations.

What the benchmarks say

DeepSeek-reported (think-max) plus independent evaluations
BenchmarkV4 ProV4 FlashContext
LiveCodeBench93.591.6Best in DeepSeek’s comparison table, above Gemini-3.1-Pro
Codeforces rating32063052Above GPT-5.4 (3168)
SWE-bench Verified80.679.0Claude Opus 4.6: 80.8 — effectively tied
GPQA Diamond90.188.1Gemini-3.1-Pro: 94.3
MMLU-Pro87.586.2Ties GPT-5.4
AA Intelligence Index (independent)5247#2 open-weights family behind Kimi K2.6; Flash ≈ Sonnet-4.6 class
SWE-bench Verified (NIST CAISI, independent)74%—GPT-5.5: 81% — CAISI places V4 ~8 months behind the US frontier

Self-reported numbers from the V4 Pro model card; independent rows from Artificial Analysis and the NIST CAISI evaluation. AA flags heavy verbosity and weak factual recall (its Omniscience test) — worth knowing for knowledge-heavy workloads.

The pricing story

V4 launched with Pro at a list price of $1.74/$3.48 per million tokens, immediately “discounted 75%” to $0.435/$0.87 — a discount that quietly became the standard price by June. Flash sits at $0.14/$0.28, undercutting even DeepSeek’s previous V3.2 while multiplying context 8×. Cache hits are absurdly cheap ($0.0028–$0.0036 per million — up to 1/120th of a miss). Next up, per TechNode: the full V4 release introduces DeepSeek’s first peak/off-peak pricing, with peak hours billed at double the off-peak rate. Smouter’s live rate for deepseek-v4-pro right now: $0.217 in / $0.435 out — 50% under DeepSeek’s standard price, all hours, no peak windows.

Adoption, and the fine print

By July 2026, V4 Flash alone moved more than double the weekly tokens of the most-used frontier model on OpenRouter — while costing ~23× less per token on average. The fine print from three months of production use: recurring peak-time 429/503 “server busy” responses on the first-party API, dynamic rate limits you cannot buy your way past, and hands-on tests (like Kilo’s build-a-project rubric) that found it capable but imperfect — Pro scored between Claude Opus and Kimi, and Flash built a near-working project for two cents.

Start routing DeepSeek V4 Pro today.Free to start · no card · one key for every model
Get started free