Qwen3 Coder Next, explained
Facts checked 2026-07-21Qwen3-Coder-Next is Alibaba’s open-weight coding-agent model, released on 2026-02-02 under Apache 2.0. It activates just 3 billion of its 80 billion parameters per token, yet lands SWE-bench numbers that took 10–20× more active compute a year earlier — which is why it became the default budget engine inside coding agents like Cline within days of launch.
The Qwen machine
Qwen is Alibaba Cloud’s model family, open-sourced since September 2023 and shipped at every size point under permissive licenses. The strategy produced the most-derived open family in the world: over 200,000 Qwen variants on Hugging Face by March 2026 (Wikipedia). Behind it sits real money — a RMB 380 billion (~$53B) three-year AI infrastructure plan that Alibaba’s CEO has since said the company will exceed (CNBC).
A footnote with weight: within a month of this model shipping, the Qwen team lost three senior leaders — including coding lead Binyuan Hui, whose group built the Coder line, and overall tech lead Junyang Lin, who signed off with “bye my beloved qwen.” Alibaba responded with a restructuring and an internal AGI task force (CNBC). Coder-Next is effectively the last coder flagship of that era.
From 480B to 80B
The line runs: Qwen2.5-Coder (November 2024) → Qwen3 (April 2025) → Qwen3-Coder-480B-A35B (July 2025, the first big agentic coder, which Cerebras later served at up to 2,000 tokens/second) → Qwen3-Next-80B (September 2025, a new hybrid architecture) → Qwen3-Coder-Next (February 2026), which puts the coder recipe on the efficient Next architecture. It supports 358 programming languages and a 262,144-token native context (extendable to 1M with YaRN).
Why 3B active parameters is enough
- Hybrid attention: 48 layers alternating 36 linear-attention (Gated DeltaNet) layers with 12 gated full-attention layers — the trick that keeps quarter-million-token agent loops cheap.
- Ultra-sparse MoE: 512 experts, 10 + 1 shared active per token → 80B total, ~3B active.
- Non-thinking by design: it never emits think-blocks, which Qwen pitches as simpler and faster for production agents.
- Trained like an agent: tasks synthesized from real GitHub PRs, run in executable environments across six scaffolds (SWE-agent, OpenHands, Claude Code, Qwen Code and more), with a 480B teacher distilling domain experts back into one model.
- Local-first: ~46GB at 4-bit — the first agent-grade coder that fits a 64GB MacBook Pro, a point Simon Willison made in the 735-point Hacker News launch thread.
What the benchmarks say
| Benchmark | Qwen3-Coder-Next | Context |
|---|---|---|
| SWE-bench Verified (SWE-agent) | 70.6% | Claude Sonnet 4.5: 76.0 · DeepSeek-V3.2: 70.2 |
| SWE-bench Pro | 44.3% | Equals Sonnet 4.5 (44.3), best open model shown |
| Aider code editing | 66.2 | DeepSeek-V3.2: 69.9 · GLM-4.7: 52.1 |
| LiveCodeBench v6 | 58.9 | vs its own 480B sibling: 44.9 |
| Terminal-Bench 2.0 | 36.2% | Claude Opus 4.5: 57.3 — terminals are the honest gap |
| AA Intelligence Index v4.1 (independent) | 21 — #2 of 39 in class | Class average: 7 |
Sources: the announcement, technical report, and Artificial Analysis. AA also measures ~98 tok/s and notes it is one of the most verbose models tested — the flip side of cheap active compute.
Pricing: tiers vs flat
Alibaba’s own Model Studio bills Coder-Next in tiers by request size: $0.30/$1.50 per million tokens up to 32K input, $0.50/$2.50 to 128K, $0.80/$4.00 to 256K — and one big-context request bills *all* its tokens at the tier it lands in. Third-party hosting undercuts that (the cheapest hosts sit around $0.11/$0.80 after ninety days of price war), and adoption ran ahead of it: 1.58 billion prompt tokens through OpenRouter on day one, with coding agents pi, Hermes Agent and Cline its biggest public consumers. Smouter’s live rate right now: $0.081 in / $0.46 out flat — 71% under Alibaba’s lowest tier, whatever the request size.
Sources
- Qwen — Qwen3-Coder-Next announcement
- Hugging Face — Qwen/Qwen3-Coder-Next model card (specs, license, scaffolds)
- Qwen3-Coder-Next technical report (benchmarks, RL details, reward hacking)
- Alibaba Cloud Model Studio — official tiered pricing
- Artificial Analysis — Qwen3-Coder-Next (index, speed, verbosity)
- OpenRouter — qwen3-coder-next (provider prices, top apps)
- Hacker News — launch thread (2026-02-03)
- Unsloth — running Coder-Next locally (46GB at 4-bit)
- CNBC — Qwen leadership exits and Alibaba restructuring (2026-03-17)
- Cerebras — Qwen3-Coder-480B at 2,000 tok/s (family speed lineage)
- Wikipedia — Qwen (family history, derivative counts)