GLM-5 — API Pricing & Benchmarks
GLM-5 is Zhipu AI (Z.ai)'s mid reasoning model, released in 2026-02. The API costs $1.00 per 1M input tokens and $3.20 per 1M output tokens ($1.55 blended at a 3:1 ratio), with cached input at $0.20. It supports a 200K-token context window. On the independently measured Artificial Analysis Intelligence Index it scores 40.6, ranking #41 of 68 models tracked on this site, with a median output speed of 48 tokens/sec.
Explore provider routes and additional benchmark sources for GLM-5
| Provider | Zhipu AI (Z.ai) |
| Input price / 1M tokens | $1.00 |
| Output price / 1M tokens | $3.20 |
| Cached input / 1M tokens | $0.20 |
| Blended price (3:1) | $1.55 |
| Context window | 200K |
| Max output tokens | 128K |
| Knowledge cutoff | — |
| License | open weights |
| Input modalities | T |
| Output speed (tok/s) | 48 |
| Released | 2026-02 |
| Artificial Analysis Intelligence Index | 40.6 |
| LMArena Elo | 1446 |
| GPQA Diamond | 82% |
| Humanity's Last Exam | 29.3% |
Where GLM-5 fits
On a blended 3:1 basis GLM-5 costs $1.55 per 1M tokens, cheaper than 55% of the 80 priced models in this index. It sits in the mid-tier bracket where most production chat, RAG and summarization workloads run: meaningfully cheaper than flagships while keeping most of their capability. Measured output speed of 48 tokens/sec is below the index median of 104 — typical of reasoning-heavy models, and worth checking against interactive latency budgets. The weights are open, so it can be self-hosted for data-residency or cost reasons instead of consumed via API. Cached input is priced at $0.20 (80% below standard input), which rewards prompt structures with long stable prefixes — see our caching guide.
Price history
| Period | Input $/1M | Output $/1M | Cached $/1M |
|---|---|---|---|
| 2026-07-18 → current | $1.00 | $3.20 | $0.20 |
No list-price changes recorded since tracking began. Full dataset: price-history.json (CC BY 4.0).
Alternatives at a similar price
- Grok 4.3 (xAI) — $1.56/1M blended, AA Index 37.9
- Grok 4.20 (xAI) — $1.56/1M blended
- Gemini 3.8 Flash (Google) — $1.50/1M blended, AA Index 58.7
- Gemini 3.7 Flash (Google) — $1.50/1M blended, AA Index 56
About Zhipu AI (Z.ai)
Zhipu AI (Z.ai) publishes the GLM series as open weights. GLM-5.3 benchmarks within a few points of the proprietary frontier at roughly a third of the price, and GLM-5.3-Flash keeps most of that score at small-model pricing.