Skip to content
Groq / Together listSmouter (cached)You save
Input / 1M tokens$0.15$0.018−88%
Output / 1M tokens$0.60$0.085−86%

Smouter price is the live value-lane rate (cached) for gpt-oss-120b. Reference is the identical list price published by Groq and Together AI for gpt-oss-120b ($0.15 in / $0.60 out per 1M tokens, July 2026). Prices shown per 1M tokens.

One line to switch

Smouter speaks the OpenAI API. If your app already calls gpt-oss-120b through the vendor, OpenRouter, or the OpenAI SDK, you only change the base URL and the key — the request body is identical.

  • Drop-in chat/completions — streaming, tools, JSON mode
  • The official open weights OpenAI released — no fine-tune, no quant tricks
  • Works in any coding agent or IDE that takes an OpenAI base URL
gpt_oss_120b.py
from openai import OpenAI client = OpenAI(    base_url="https://api.smouter.ai/v1",  # the only line you change    api_key="sk-smouter-...",) resp = client.chat.completions.create(    model="gpt-oss-120b",  # open weights, tiny price    messages=[{"role": "user", "content": "Draft a release note."}],)

Frontier lab, open weights

gpt-oss-120b is OpenAI's own open-weight reasoning model — a serious model with a permissive license, which is why every host serves it and prices keep falling.

We chase the floor for you

Hosts price gpt-oss-120b everywhere from a few cents to fifteen per million input tokens. Smouter routes each request to the cheapest healthy provider, live.

Pay per token, no lock-in

No subscription, no minimums, no card to start. Add credits, get an API key, and you are routing gpt-oss-120b in minutes. Cancel any time.

gpt-oss-120b, explained

Facts checked 2026-07-21

gpt-oss-120b is OpenAI’s first open-weight language model since GPT-2 in 2019 — released 2025-08-05 under Apache 2.0, delivering o4-mini-class reasoning from a checkpoint that fits on a single 80GB GPU. A year on it remains one of the most-downloaded model families anywhere, and among the fastest-served: the efficiency pick of the open ecosystem.

Why OpenAI went open again

The pivot traces to January 2025, when DeepSeek R1 upended the closed-lab consensus. Days later, Sam Altman told a Reddit AMA that OpenAI had been on the wrong side of history on open source (Slashdot). Six months, two delays, and one adversarial-fine-tuning safety evaluation later, gpt-oss-120b and its 20B sibling shipped — two weeks after a US AI Action Plan that explicitly urged American labs to release open models. Altman framed the launch as an open AI stack “created in the United States” (TechCrunch).

Inside the model

  • 116.8B total parameters, ~5.1B active — 36 layers, 128 experts per block with top-4 routing, grouped-query attention with learned attention sinks, 131,072-token context.
  • MXFP4 quantization brings the checkpoint to 60.8 GiB — one H100 or MI300X runs it; community MoE-offload recipes run it on an 8GB-VRAM GPU with 64GB of system RAM.
  • Adjustable reasoning effort (low/medium/high) — one system-message line moves AIME 2024 from 56.3% to 95.8%.
  • Fully exposed chain-of-thought, deliberately left unsupervised in training so developers can monitor it — plus the open-sourced “harmony” prompt format the model requires.
  • Training ran 2.1M H100-hours — an estimated $4.2M–$23.1M at 2025 rental rates. Knowledge cutoff June 2024.

What the benchmarks say

OpenAI model card (high effort) vs OpenAI’s own closed models
Benchmarkgpt-oss-120bo3o4-mini
AIME 2024 (no tools)95.891.693.4
GPQA Diamond80.183.381.4
MMLU90.0——
Codeforces Elo (tools)262227062719
SWE-bench Verified62.469.168.1
HealthBench57.6≈60≈54

From the official model card. Today’s placement per Artificial Analysis: #5 of 62 open models in its size class for intelligence, #3 for speed (~312 tok/s median across hosts). Newer, far larger open models (GLM-5.2, DeepSeek V4) have passed it on raw intelligence — nothing near its size has.

The honest catch

OpenAI’s own card reports high hallucination rates on people-and-facts tests — it fabricated on 49% of PersonQA and 78% of SimpleQA attempts, roughly triple o1’s rate. The community’s “benchmaxxed” critique — stellar STEM scores, thinner world knowledge and human preference (LMArena rank sits far below its benchmark tier) — matches its STEM-heavy training data. Use it where reasoning density matters and facts are grounded by your context or tools; pair it with retrieval for knowledge work.

No first-party API — the market sets the price

OpenAI never put gpt-oss in its own API, so pricing is a pure hosting market: Groq and Together list $0.15/$0.60 per million tokens, the aggregate floor runs as low as $0.03/$0.18, and speed specialists charge a premium — Cerebras set a 3,000 tokens/second world record on launch day. Smouter’s live rate right now: $0.018 in / $0.085 out — 87% under the Groq/Together list. Adoption is still compounding a year in: ~4.4M Hugging Face downloads in the last 30 days and 11M+ on Ollama.

Favorite footnotes: in OpenAI’s own launch demo the model was asked how many experts per layer it has, didn’t know, and browsed the web (28 tool calls) to find its own leaked specs. Researchers extracted an approximation of the un-aligned base model from its 20B sibling within about a week — a live lesson in what “open weights” irreversibly means.

Start routing gpt-oss-120b today.Free to start · no card · one key for every model
Get started free