Skip to content
Moonshot listSmouterYou save
Input / 1M tokens$3.00$2.16−28%
Output / 1M tokens$15.00$10.78−28%

Smouter price is the live value-lane rate (cached) for kimi-k3. Reference is Moonshot's published API list price ($3 in / $15 out per 1M tokens, July 2026). Prices shown per 1M tokens.

One line to switch

Smouter speaks the OpenAI API. If your app already calls Kimi through OpenRouter, Moonshot, or the OpenAI SDK, you only change the base URL and the key — the request body is identical.

  • Drop-in chat/completions — streaming, tools, JSON mode
  • Full 1M-token context, same weights
  • Works in any coding agent or IDE that takes an OpenAI base URL
kimi_k3.py
from openai import OpenAI client = OpenAI(    base_url="https://api.smouter.ai/v1",  # the only line you change    api_key="sk-smouter-...",) resp = client.chat.completions.create(    model="kimi-k3",  # same model, below list price    messages=[{"role": "user", "content": "Refactor this function."}],)

Same model, same code

Kimi K3 exactly as Moonshot ships it — full 1M-token context. Point your existing OpenAI SDK at Smouter and change one line. No rewrite, no new client.

One key, every model

The same key that serves Kimi K3 also serves GPT, Claude, Gemini, DeepSeek and 50+ others. Switch models with a string, not a new integration.

Pay per token, no lock-in

No subscription, no minimums, no card to start. Add credits, get an API key, and you are routing Kimi K3 in minutes. Cancel any time.

Kimi K3, explained

Facts checked 2026-07-21

Kimi K3 is Moonshot AI’s frontier model, announced on 2026-07-16 — a 2.8-trillion-parameter, natively multimodal model with a million-token context window. Within days of launch it became the highest-scoring open-weights-track model ever listed on Artificial Analysis’ Intelligence Index and took the #1 spot on LMArena’s WebDev arena. Here is where it came from, what is inside it, and what the numbers actually say.

The company behind it

Moonshot AI was founded in March 2023 by three Tsinghua University schoolmates — Yang Zhilin, Zhou Xinyu and Wu Yuxin — and is named after Pink Floyd’s *The Dark Side of the Moon*. It is one of China’s so-called “Six AI Tigers,” backed by Alibaba and Tencent, and was last reported at a $3.8B valuation with a Hong Kong IPO under consideration (Wikipedia).

Its consumer app, Kimi, made long context a household feature in China — launching in October 2023 with the ability to digest roughly 200,000 Chinese characters in one conversation, and going so viral by March 2024 that a two-day outage forced a public apology. The serving platform built for it, Mooncake, processes on the order of 100 billion tokens a day and won a Best Paper award at USENIX FAST.

The road to K3

K3 caps a remarkable twelve months. In July 2025, Kimi K2 became the first trillion-parameter open-weight model — the most-downloaded model on Hugging Face the day after release, and the moment *Nature* called “another DeepSeek moment.” Moonshot then shipped relentlessly: K2 Thinking (November 2025, an open reasoning model that could chain 200–300 tool calls), K2.5 (January 2026, native vision), K2.6 (April 2026, coding), and K2.7 Code (June 2026). Each major release pulled around a million Hugging Face downloads.

Moonshot’s own launch claim is that Kimi models have set the upper bound of open-model scale in nine of the twelve months from July 2025 to July 2026 (Kimi K3 announcement). K3 itself launched proprietary-first, with full weights pledged for 2026-07-27 and a technical report to follow; every K2-family release used a Modified MIT license, though K3’s license is not yet announced.

Inside the model

  • 2.8 trillion total parameters — the first open model at that scale. Active-parameter count is not yet disclosed; the MoE routes 16 of 896 experts per token.
  • Kimi Delta Attention (KDA) — a hybrid linear-attention design — plus “Attention Residuals,” which retrieve representations across depth instead of accumulating them.
  • 1,048,576-token context window with native image input (video understanding in the Kimi apps).
  • Trained with Per-Head Muon, an evolution of the Muon optimizer Moonshot scaled for K2, with quantization-aware training from the SFT stage — weights ship in MXFP4.
  • Moonshot claims the combined changes deliver ~2.5× the scaling efficiency of K2. Training cost and token counts are undisclosed.

What the benchmarks say

The honest framing — which is also Moonshot’s own — is that K3 still trails Claude Fable 5 and GPT-5.6 Sol overall, and beats essentially everything else it was tested against, including Claude Opus 4.8, GPT-5.5 and GLM-5.2. Five days in, most numbers are vendor-reported; the independent signals so far are Artificial Analysis and LMArena.

Selected results (Moonshot-reported at max reasoning unless noted)
BenchmarkKimi K3Best closed comparison
Artificial Analysis Intelligence Index v4.1 (independent)57 — #4 overall, #1 open-trackClaude Fable 5: 59.9
LMArena WebDev arena (independent)#1, Elo 1677—
LMArena text arena (independent)#8, ~1487 Elo—
Terminal-Bench 2.188.3GPT-5.6 Sol: 88.8
GPQA-Diamond93.5GPT-5.6 Sol: 94.1
BrowseComp (with compaction)91.2GPT-5.6 Sol: 90.4
DeepSWE v1.1 (agentic SWE)67.5GPT-5.6 Sol: 73.0
Humanity’s Last Exam (no tools)43.5Claude Fable 5: 53.3

Full table with sources in the K3 announcement; comparisons mix Moonshot’s runs with published leaderboard numbers. Classic suites like SWE-bench Verified and AIME have no published K3 numbers yet.

Two practical caveats from Artificial Analysis’ independent measurement: served from Kimi’s own API, K3 output ran at ~39.5 tokens/second, and the model is, in their words, “notably slow and very verbose” — verbosity that cost $2,709.75 to run their full index eval. Fast third-party hosting will likely follow the weights release.

Pricing: the end of the shock-cheap era

Moonshot’s list price for K3 is $3.00 per million input tokens ($0.30 on cache hits) and $15.00 per million output tokens. That is a deliberate 3–5× step up from the K2 era — K2 launched at $0.15/$2.50 in 2025, and K2.6 still lists at $0.95/$4.00 — pricing K3 at Claude-Sonnet-class rates while claiming near-flagship intelligence. Moonshot counters that its coding workloads see >90% cache-hit rates, making effective cost far lower than list. Smouter’s live rate for kimi-k3 right now: $2.16 in / $10.78 out — 28% under Moonshot’s list, same weights, same 1M context.

The week it moved markets

K3’s launch coincided with a Xi Jinping speech at the World AI Conference and knocked the Nasdaq down about 1% the next day (TechCrunch). Pre-IPO markets repriced the closed labs: IG measured a combined ~$392B drop in Anthropic’s and OpenAI’s implied valuations in the five days after launch. In Washington, reporting by Axios describes a revived push to restrict Chinese models — alongside the admission that banning downloadable weights is nearly unenforceable. Meanwhile adoption keeps compounding: Coinbase’s CEO says running GLM and Kimi models in production roughly halved the company’s AI spend.

Start routing Kimi K3 today.Free to start · no card · one key for every model
Get started free