gpt-oss-120b, explained
Facts checked 2026-07-21gpt-oss-120b is OpenAI’s first open-weight language model since GPT-2 in 2019 — released 2025-08-05 under Apache 2.0, delivering o4-mini-class reasoning from a checkpoint that fits on a single 80GB GPU. A year on it remains one of the most-downloaded model families anywhere, and among the fastest-served: the efficiency pick of the open ecosystem.
Why OpenAI went open again
The pivot traces to January 2025, when DeepSeek R1 upended the closed-lab consensus. Days later, Sam Altman told a Reddit AMA that OpenAI had been on the wrong side of history on open source (Slashdot). Six months, two delays, and one adversarial-fine-tuning safety evaluation later, gpt-oss-120b and its 20B sibling shipped — two weeks after a US AI Action Plan that explicitly urged American labs to release open models. Altman framed the launch as an open AI stack “created in the United States” (TechCrunch).
Inside the model
- 116.8B total parameters, ~5.1B active — 36 layers, 128 experts per block with top-4 routing, grouped-query attention with learned attention sinks, 131,072-token context.
- MXFP4 quantization brings the checkpoint to 60.8 GiB — one H100 or MI300X runs it; community MoE-offload recipes run it on an 8GB-VRAM GPU with 64GB of system RAM.
- Adjustable reasoning effort (low/medium/high) — one system-message line moves AIME 2024 from 56.3% to 95.8%.
- Fully exposed chain-of-thought, deliberately left unsupervised in training so developers can monitor it — plus the open-sourced “harmony” prompt format the model requires.
- Training ran 2.1M H100-hours — an estimated $4.2M–$23.1M at 2025 rental rates. Knowledge cutoff June 2024.
What the benchmarks say
| Benchmark | gpt-oss-120b | o3 | o4-mini |
|---|---|---|---|
| AIME 2024 (no tools) | 95.8 | 91.6 | 93.4 |
| GPQA Diamond | 80.1 | 83.3 | 81.4 |
| MMLU | 90.0 | — | — |
| Codeforces Elo (tools) | 2622 | 2706 | 2719 |
| SWE-bench Verified | 62.4 | 69.1 | 68.1 |
| HealthBench | 57.6 | ≈60 | ≈54 |
From the official model card. Today’s placement per Artificial Analysis: #5 of 62 open models in its size class for intelligence, #3 for speed (~312 tok/s median across hosts). Newer, far larger open models (GLM-5.2, DeepSeek V4) have passed it on raw intelligence — nothing near its size has.
The honest catch
OpenAI’s own card reports high hallucination rates on people-and-facts tests — it fabricated on 49% of PersonQA and 78% of SimpleQA attempts, roughly triple o1’s rate. The community’s “benchmaxxed” critique — stellar STEM scores, thinner world knowledge and human preference (LMArena rank sits far below its benchmark tier) — matches its STEM-heavy training data. Use it where reasoning density matters and facts are grounded by your context or tools; pair it with retrieval for knowledge work.
No first-party API — the market sets the price
OpenAI never put gpt-oss in its own API, so pricing is a pure hosting market: Groq and Together list $0.15/$0.60 per million tokens, the aggregate floor runs as low as $0.03/$0.18, and speed specialists charge a premium — Cerebras set a 3,000 tokens/second world record on launch day. Smouter’s live rate right now: $0.018 in / $0.085 out — 87% under the Groq/Together list. Adoption is still compounding a year in: ~4.4M Hugging Face downloads in the last 30 days and 11M+ on Ollama.
Favorite footnotes: in OpenAI’s own launch demo the model was asked how many experts per layer it has, didn’t know, and browsed the web (28 tool calls) to find its own leaked specs. Researchers extracted an approximation of the un-aligned base model from its 20B sibling within about a week — a live lesson in what “open weights” irreversibly means.
Sources
- OpenAI — Introducing gpt-oss (launch post)
- gpt-oss model card (arXiv:2508.10925) — architecture, benchmarks, hallucination data
- Artificial Analysis — gpt-oss-120b today (class rank, speed)
- OpenRouter — live provider pricing (~18 hosts)
- TechCrunch — launch coverage (2025-08-05)
- Slashdot — Altman’s “wrong side of history” AMA (2025-01-31)
- Cerebras — 3,000 tok/s world record on gpt-oss-120b
- Simon Willison — day-one hands-on (speeds, local RAM, cost estimate)
- Sebastian Raschka — from GPT-2 to gpt-oss architecture teardown
- Hugging Face — openai/gpt-oss-120b (downloads, ecosystem)