Skip to content
Alibaba listSmouter (cached)You save
Input / 1M tokens$0.30$0.081−73%
Output / 1M tokens$1.50$0.46−69%

Smouter price is the live value-lane rate (cached) for qwen3-coder-next. Reference is Alibaba Cloud Model Studio's lowest billing tier for qwen3-coder-next ($0.30 in / $1.50 out per 1M tokens for requests up to 32K input, July 2026). Alibaba bills bigger requests at higher tiers — up to $0.80/$4.00 at 256K — while Smouter's price is one flat rate. Prices shown per 1M tokens.

One line to switch

Smouter speaks the OpenAI API. If your app already calls Qwen3 Coder Next through the vendor, OpenRouter, or the OpenAI SDK, you only change the base URL and the key — the request body is identical.

  • Drop-in chat/completions — streaming, tools, JSON mode
  • Same weights, 262K context — built for coding agents and IDE tools
  • Works in any coding agent or IDE that takes an OpenAI base URL
qwen3_coder.py
from openai import OpenAI client = OpenAI(    base_url="https://api.smouter.ai/v1",  # the only line you change    api_key="sk-smouter-...",) resp = client.chat.completions.create(    model="qwen3-coder-next",  # same model, flat per-token price    messages=[{"role": "user", "content": "Write tests for this module."}],)

Made for coding agents

Qwen3-Coder-Next is tuned for tool use and multi-step coding work, and Smouter speaks the OpenAI API — so it drops into any agent or IDE that takes a base URL.

Flat price, no tiers

Alibaba bills big-context requests at higher per-token tiers. Smouter's price is one flat live rate per token, whatever the request size.

The rest of Qwen too

Qwen3.7-Max, 3.7-Plus and a dozen smaller Qwen3 variants are on the same key — pick per request with a one-string change.

Qwen3 Coder Next, explained

Facts checked 2026-07-21

Qwen3-Coder-Next is Alibaba’s open-weight coding-agent model, released on 2026-02-02 under Apache 2.0. It activates just 3 billion of its 80 billion parameters per token, yet lands SWE-bench numbers that took 10–20× more active compute a year earlier — which is why it became the default budget engine inside coding agents like Cline within days of launch.

The Qwen machine

Qwen is Alibaba Cloud’s model family, open-sourced since September 2023 and shipped at every size point under permissive licenses. The strategy produced the most-derived open family in the world: over 200,000 Qwen variants on Hugging Face by March 2026 (Wikipedia). Behind it sits real money — a RMB 380 billion (~$53B) three-year AI infrastructure plan that Alibaba’s CEO has since said the company will exceed (CNBC).

A footnote with weight: within a month of this model shipping, the Qwen team lost three senior leaders — including coding lead Binyuan Hui, whose group built the Coder line, and overall tech lead Junyang Lin, who signed off with “bye my beloved qwen.” Alibaba responded with a restructuring and an internal AGI task force (CNBC). Coder-Next is effectively the last coder flagship of that era.

From 480B to 80B

The line runs: Qwen2.5-Coder (November 2024) → Qwen3 (April 2025) → Qwen3-Coder-480B-A35B (July 2025, the first big agentic coder, which Cerebras later served at up to 2,000 tokens/second) → Qwen3-Next-80B (September 2025, a new hybrid architecture) → Qwen3-Coder-Next (February 2026), which puts the coder recipe on the efficient Next architecture. It supports 358 programming languages and a 262,144-token native context (extendable to 1M with YaRN).

Why 3B active parameters is enough

  • Hybrid attention: 48 layers alternating 36 linear-attention (Gated DeltaNet) layers with 12 gated full-attention layers — the trick that keeps quarter-million-token agent loops cheap.
  • Ultra-sparse MoE: 512 experts, 10 + 1 shared active per token → 80B total, ~3B active.
  • Non-thinking by design: it never emits think-blocks, which Qwen pitches as simpler and faster for production agents.
  • Trained like an agent: tasks synthesized from real GitHub PRs, run in executable environments across six scaffolds (SWE-agent, OpenHands, Claude Code, Qwen Code and more), with a 480B teacher distilling domain experts back into one model.
  • Local-first: ~46GB at 4-bit — the first agent-grade coder that fits a 64GB MacBook Pro, a point Simon Willison made in the 735-point Hacker News launch thread.

What the benchmarks say

Qwen-reported unless noted (max 300 agent turns)
BenchmarkQwen3-Coder-NextContext
SWE-bench Verified (SWE-agent)70.6%Claude Sonnet 4.5: 76.0 · DeepSeek-V3.2: 70.2
SWE-bench Pro44.3%Equals Sonnet 4.5 (44.3), best open model shown
Aider code editing66.2DeepSeek-V3.2: 69.9 · GLM-4.7: 52.1
LiveCodeBench v658.9vs its own 480B sibling: 44.9
Terminal-Bench 2.036.2%Claude Opus 4.5: 57.3 — terminals are the honest gap
AA Intelligence Index v4.1 (independent)21 — #2 of 39 in classClass average: 7

Sources: the announcement, technical report, and Artificial Analysis. AA also measures ~98 tok/s and notes it is one of the most verbose models tested — the flip side of cheap active compute.

Pricing: tiers vs flat

Alibaba’s own Model Studio bills Coder-Next in tiers by request size: $0.30/$1.50 per million tokens up to 32K input, $0.50/$2.50 to 128K, $0.80/$4.00 to 256K — and one big-context request bills *all* its tokens at the tier it lands in. Third-party hosting undercuts that (the cheapest hosts sit around $0.11/$0.80 after ninety days of price war), and adoption ran ahead of it: 1.58 billion prompt tokens through OpenRouter on day one, with coding agents pi, Hermes Agent and Cline its biggest public consumers. Smouter’s live rate right now: $0.081 in / $0.46 out flat — 71% under Alibaba’s lowest tier, whatever the request size.

Start routing Qwen3 Coder Next today.Free to start · no card · one key for every model
Get started free