GLM-4.5-Air — API Pricing & Benchmarks
GLM-4.5-Air is Zhipu AI (Z.ai)'s small reasoning model, released in 2025-07. The API costs $0.20 per 1M input tokens and $1.10 per 1M output tokens ($0.43 blended at a 3:1 ratio), with cached input at $0.03. On the independently measured Artificial Analysis Intelligence Index it scores 16.7, ranking #59 of 68 models tracked on this site, with a median output speed of 81 tokens/sec.
Explore provider routes and additional benchmark sources for GLM-4.5-Air
| Provider | Zhipu AI (Z.ai) |
| Input price / 1M tokens | $0.20 |
| Output price / 1M tokens | $1.10 |
| Cached input / 1M tokens | $0.03 |
| Blended price (3:1) | $0.43 |
| Context window | — |
| Max output tokens | 96K |
| Knowledge cutoff | — |
| License | open weights |
| Input modalities | T |
| Output speed (tok/s) | 81 |
| Released | 2025-07 |
| Artificial Analysis Intelligence Index | 16.7 |
| LMArena Elo | 1384 |
| GPQA Diamond | 73.3% |
| Humanity's Last Exam | 7% |
Where GLM-4.5-Air fits
On a blended 3:1 basis GLM-4.5-Air costs $0.43 per 1M tokens, cheaper than 86% of the 80 priced models in this index. It is built for high-volume, latency-sensitive work — classification, extraction, routing and autocomplete — where per-token price matters more than peak intelligence. The weights are open, so it can be self-hosted for data-residency or cost reasons instead of consumed via API. Cached input is priced at $0.03 (85% below standard input), which rewards prompt structures with long stable prefixes — see our caching guide.
Price history
| Period | Input $/1M | Output $/1M | Cached $/1M |
|---|---|---|---|
| 2026-07-18 → current | $0.20 | $1.10 | $0.03 |
No list-price changes recorded since tracking began. Full dataset: price-history.json (CC BY 4.0).
Alternatives at a similar price
- Codestral 25.08 (Mistral AI) — $0.45/1M blended
- GPT-5.6 Luna (OpenAI) — $0.45/1M blended, AA Index 52.3
- GPT-5.4 nano (OpenAI) — $0.46/1M blended, AA Index 39.7
- K-EXAONE 236B (LG AI Research) — $0.35/1M blended
About Zhipu AI (Z.ai)
Zhipu AI (Z.ai) publishes the GLM series as open weights. GLM-5.3 benchmarks within a few points of the proprietary frontier at roughly a third of the price, and GLM-5.3-Flash keeps most of that score at small-model pricing.