LLM API Price Changes — August 2026
August 9, 2026 · AI API Index
This is the monthly change log for LLM API list prices. I check every number here against the provider's own pricing page before it goes in, and the full history is published as price-history.json (CC BY 4.0) if you want the raw data.
The big one: GPT-5.6 Luna, down 80%
On August 8, OpenAI dropped GPT-5.6 Luna from $1.00 input / $6.00 output to $0.20 / $1.20 per million tokens. Cached input went from $0.10 to $0.02. I checked this one twice, because an 80% cut on a model that shipped in July is not a normal event. It's real.
Some context on why it matters. Luna now costs exactly what the previous-generation GPT-5.4 nano costs ($0.20 / $1.25), and nano was already the bargain option. On the Artificial Analysis index Luna scores 52.3 against nano's 39.7. It also undercuts Gemini 3.5 Flash-Lite ($0.30 / $2.50) while scoring well above it. If you run any high-volume pipeline on a small model, spend a day re-running your evals with Luna in the mix. At these prices the switch could pay for the eval time in a week.
Terra gets a quieter trim
GPT-5.6 Terra, the mid-tier sibling, moved from $2.50 / $15.00 to $2.00 / $12.00, with cached input down to $0.20. That's 20%. Terra now matches Claude Sonnet 5 on input price, though Sonnet still writes cheaper at $10 per million output tokens. Percentage-wise this looks tame next to Luna, but the mid-tier is where most production chat and agent traffic actually lives, so in absolute dollars I suspect this cut saves buyers more money than the flashy one.
Qwen3.8-Max lands in the top five
Alibaba released Qwen3.8-Max this month and it entered the index at $2.00 / $6.00 with an AA score of 58.1. Fifth overall as of this writing, behind Claude Opus 5, Claude Fable 5, GPT-5.6 Sol and Kimi K3. Blended at the usual 3:1 ratio it works out to about $3 per million tokens; GPT-5.6 Sol blends to a bit over $11. Whether the score gap between them matters for your workload is something only your own eval can answer. The price gap is 3.7×. That's why this is the model I'd test first this month. See the head-to-head pages for detailed comparisons.
Smaller notes
- K-EXAONE 236B (LG AI Research) corrected to $0.20 / $0.80, cached at $0.10. Among the cheapest open-weights flagships I track.
- Qwen3.8-Max's cached-input price settled at $0.25 after an initial $0.20 listing.
- Amazon Nova 2 Lite now lists cached input at $0.075.
If you're buying
Two things stood out to me this month. First, the floor keeps dropping. The cheapest model that can reliably follow instructions got 80% cheaper in a single announcement, so any unit-economics spreadsheet from July is already stale. Second, the pressure is coming from open-weights and Chinese labs. Qwen3.8-Max at $2/$6 and Kimi K3 at $3/$15 sit right on the price points OpenAI and Anthropic charge for mid-tier models, with scores close to flagship level. So far the incumbents have answered with targeted cuts on their smaller models and left flagship prices alone. Opus 5 and Sol haven't moved a cent.
To see what any of this does to your own bill, drop your monthly token volumes into the cost calculator. It uses current list and cached prices and sorts by what you'd actually pay.