GLM-5.2, explained
Facts checked 2026-07-21GLM-5.2 is Z.ai’s open-weight frontier model — 744B parameters (40B active), a million-token context, MIT license — released in mid-June 2026. It is currently the highest-scoring open-weight model on Artificial Analysis’ Intelligence Index, the model a former White House AI czar called “just a tick below Opus 4.8,” and the one Interconnects dubbed the step change for open models. It is also the single most-routed model on Smouter.
From a Tsinghua lab to the Hong Kong exchange
Zhipu AI spun out of Tsinghua University’s Knowledge Engineering Group in 2019; the GLM name comes from its 2021 “General Language Model” pretraining research. It open-sourced GLM-130B back in 2022, was added to the US Entity List in January 2025, rebranded internationally as Z.ai, and then did something no other frontier lab had done: on 2026-01-08 it went public in Hong Kong, raising ~$558M as China’s first major listed LLM developer (Wikipedia). After GLM-5.2 shipped, the stock at one point traded around 13× its IPO price, with JPMorgan projecting revenue growth above 500% for 2026.
The road to 5.2
The modern run started with GLM-4.5 (July 2025, MIT, 355B/32B) — the release that also birthed the viral “$3-a-month coding plan.” GLM-5 (February 2026, one month after the IPO) scaled to 744B/40B over 28.5T tokens and adopted DeepSeek-style sparse attention. GLM-5.1 (April 2026) jumped hard on agentic coding. GLM-5.2 rolled out in a staggered June 2026 launch: subscribers got it June 13, and the open weights, benchmark table, and pay-per-token API followed on June 16 — day-one on Cloudflare Workers AI and, within weeks, 28 hosting providers.
Inside the model
- 744B total / 40B active MoE — 78 layers, 256 routed experts + 1 shared, 8 active per token, latent (MLA-style) attention.
- “Solid 1M” context: the window grew from 200K to 1,048,576 tokens, trained specifically on long coding-agent trajectories. IndexShare reuses one sparse-attention indexer across every four layers — 2.9× lower per-token compute at 1M context.
- An improved multi-token-prediction layer buys up to +20% speculative-decoding acceptance — cheaper, faster serving.
- Post-trained with Z.ai’s open-sourced slime RL framework, including an anti-reward-hacking module that intercepts cheating tool calls mid-rollout.
- Thinking-effort levels (High/Max) — community consensus is to always run Max. It is verbose: Artificial Analysis measured 140M tokens to finish its eval suite, versus a 96M average.
What the benchmarks say
| Benchmark | GLM-5.2 | Comparison |
|---|---|---|
| SWE-bench Pro | 62.1 | Claude Opus 4.8: 69.2 · GPT-5.5: 58.6 |
| Terminal-Bench 2.1 (best harness) | 82.7 | Claude Opus 4.8: 78.9 · GPT-5.5: 83.4 |
| AIME 2026 | 99.2 | GPT-5.5: 98.3 |
| GPQA-Diamond | 91.2 | Opus 4.8 / GPT-5.5: 93.6 |
| AA Intelligence Index v4.1 (independent) | 51 — top open-weight model | Kimi K3: 57 · Gemini 3.5 Flash: 50.2 |
| LMArena WebDev arena (independent) | #4 overall — top open model | Behind Kimi K3, Claude Fable 5, GPT-5.6-Sol |
| NIST CAISI assessment (independent) | ≈ GPT-5.2 overall; ≈ Opus 4.6 on cyber | “Probably the most capable open-weight model” at release |
Sources: the GLM-5.2 model card, Artificial Analysis, and the NIST CAISI assessment, which also notes its safety guardrails are looser than US reference models — worth knowing if you expose it directly to end users.
The step-change moment
Nathan Lambert’s widely-cited Interconnects analysis put numbers on the vibe: GLM-5.2 landed 204 days behind the comparable closed frontier (Claude Opus 4.5 → GLM-5.2), and it is the first open model that genuinely holds up as an everyday agent inside coding harnesses — a reception he compared only to DeepSeek R1’s moment. The launch timing amplified it: 5.2 arrived during a stretch when Anthropic’s newest model was unavailable to many developers, and switching was one config line away. Real-world signal: Coinbase’s CEO says running GLM and Kimi models in production roughly halved the company’s AI bill.
Pricing
Z.ai’s list price for 5.2 is $1.40 per million input tokens ($0.26 cached) and $4.40 output — about double the GLM-4.x era, but still roughly a sixth of GPT-5.5’s list and a tenth of Claude Fable 5’s. Because the weights are MIT, the hosting market undercuts list immediately: OpenRouter’s blended rate across 28 providers ran $0.84 in / $2.64 out in late July 2026, and the famous $3/month coding subscription now starts at $18/month with 5.2 billed at a quota premium. Smouter’s live rate for glm-5.2 right now: $0.47 in / $1.48 out — 66% under Z.ai’s list, and below the blended market rate above.
Sources
- Z.ai — GLM-5.2 release blog (IndexShare, 1M context, RL details)
- Hugging Face — zai-org/GLM-5.2 model card (MIT, benchmarks, methodology)
- Z.ai docs — official API pricing
- Artificial Analysis — GLM-5.2 (top open-weight, speed, verbosity)
- NIST CAISI — assessment of Z.ai’s GLM-5.2 (2026-07-17)
- Interconnects (Nathan Lambert) — “GLM-5.2 is the step change for open models”
- OpenRouter — z-ai/glm-5.2 (28 providers, blended pricing)
- GLM-5 technical report — “From Vibe Coding to Agentic Engineering”
- Wikipedia — Z.ai (history, Entity List, Hong Kong IPO)
- Memeburn — GLM-5.2 explained (market reaction, pricing comparisons)
- Web Archive — the original $3/month GLM Coding Plan (2025-09)