LLM API Price Changes — September 2026
September 2, 2026 · AI API Index
Updated September 5: three more launches landed the day after I published this, so there's a new section at the bottom covering GPT-6 Astra, Gemini 3.8 Flash and Muse Spark 1.3.
Monthly change log for LLM API list prices, part two. Same rules as last time: every number below was checked against the provider's own pricing page before it went into the index, and the raw history is in price-history.json (CC BY 4.0). This one covers August 9 through September 2, which turned out to be a busy stretch.
Claude Fable 5.1: same list price, a much cheaper cache
Anthropic shipped Claude Fable 5.1 on September 1. Input and output stay at $10 / $50 per million tokens, exactly what Fable 5 charged. The change is in cache reads, which dropped from $1.00 to $0.25. That's 2.5% of the input price. Every other Claude model (and most of the market) charges 10%.
How much that matters depends entirely on how repetitive your prompts are. Anthropic's own estimate is roughly 25% off a typical workload and up to 45% off a heavily agentic one, where the same long context gets re-read on every tool call. I'd treat those as upper bounds until you've measured your own cache hit rate, but the direction is clear. On the Artificial Analysis index it scores 65.7, the highest number that leaderboard has recorded so far, ahead of Opus 5 at 63.1. Fable 5 is still sold, now as a legacy model, with its old $1.00 cache-read price. There's no reason I can see to stay on it.
One warning for API users: 5.1 is not a drop-in swap. Forced tool choice now returns an error, and thinking blocks are tied to the model that produced them (you can't replay them elsewhere, and editing earlier turns of a conversation invalidates them). Read the migration guide before you flip the model string.
GPT-5.6 Sol: 20% off input, 33% off output, until November 21
On August 21 OpenAI cut GPT-5.6 Sol from $5.00 / $30.00 to $4.00 / $20.00, with cached input down from $0.50 to $0.40. It's labelled promotional and guaranteed through at least November 21, 2026. Amazon Bedrock matched it the same week.
The table shows the promo rate because that's what you pay today. If you're forecasting past November, budget for both numbers. OpenAI did not say what happens on the 22nd, and I wouldn't bet on the old price coming back in full either; the whole point of the cut seems to be Opus 5 sitting at $5 / $25 with a higher index score. Terra and Luna, which got their cuts in August, didn't move.
DeepSeek now bills by the hour of the day
This is the one that will surprise people. Since 16:00 UTC on August 16, DeepSeek charges different rates during peak hours: 01:00–04:00 and 06:00–10:00 UTC, Monday through Friday. Everything else, weekends included, is off-peak.
V4 Pro (the API now serves the August 13 build) went from a flat $0.435 / $0.87 to $1.32 / $3.96 at peak and $0.66 / $1.98 off-peak. So at the worst time of day you're paying about three times more for input and four and a half times more for output than in July. V4 Flash moved from $0.14 / $0.28 to $0.44 / $1.32 peak and $0.22 / $0.66 off-peak. Cache hits rose too, from a third of a cent to $0.044 (Pro, peak).
I list the peak rate in the table, since a price table should show the most you can be charged. Two things soften this. Off-peak covers most of the week, and if your batch jobs can wait a few hours you keep half the old advantage. And the new Pro build is a much better model: vals.ai measures it at 96.4% on SWE-bench Verified, up from 80.6% for the April build, which puts it under a point behind Opus 5. Even at peak, $1.32 / $3.96 for that is still cheap. It just isn't absurd anymore.
Sonnet 5 keeps its launch price
Claude Sonnet 5 launched in June at $2 / $10, described as introductory pricing through August 31 with $3 / $15 to follow. That increase has been cancelled. $2 / $10 is now the standard rate. If you had the September bump in a forecast, take it out. (I did, and I was glad to delete the row.)
New models and their prices
- GLM-5.3 (Z.ai, August 14) lists at $1.40 / $4.40, the same as GLM-5.2, and scores 59.5 on the AA index. Seventh overall, and the weights are open.
- GLM-5.3-Flash (August 26) is the one I keep coming back to. List price $0.15 / $0.50, currently 50% off until September 9, and an index score of 57.5. That is a top-ten score at the price of a small model. The table shows the list price.
- Qwen3.8-Flash (Alibaba, August 26): $0.15 / $0.47, index 55.8, flat pricing up to 1M input tokens.
- Gemini 3.7 Flash (Google, August 13): $0.75 / $3.75 as an introductory rate through December 31, then $1.50 / $7.50. Google applied the same cut to 3.6 Flash, which had been $1.50 / $7.50.
- Grok 4.6 (xAI): same $2 / $6 as Grok 4.5. Cached input went up, from $0.30 to $0.50. Index score 60.9, tied with GPT-5.6 Sol at well under half the price.
- Solar Pro 4 (Upstage, August 10): $0.30 / $1.20, double Solar Pro 3, with a 524K context window.
- Muse Spark 1.2 (Meta): still $1.25 / $4.25. Meta now publishes a cached-input price ($0.15) and a 1M context window. There's also a "contributor" tier at $0.10 / $0.20 if you let Meta train on your prompts. The table doesn't use it.
Smaller notes
- Arena Elo scores were refreshed from the LMArena dataset published September 1. Most models moved down 10–20 points together, which is the leaderboard recalibrating, so don't read it as anyone getting worse.
- Opus 5's SWE-bench Verified score is now the vals.ai measurement (97.0%) instead of Anthropic's own 96.0%. Kimi K3 gets its first independent number there too: 93.4%.
- GPQA and Humanity's Last Exam columns now sync from Artificial Analysis for every model AA measures, so a few values shifted by a point or two.
If you're buying
Three weeks ago the story was cheap models getting cheaper. This month is more mixed. DeepSeek raised prices, which I don't think has happened before. OpenAI went the other way on its flagship, and Anthropic left list prices alone but cut the part of the bill agent builders feel most. The floor is still falling, though. If you'd told me in July that a 57-point model would list at fifteen cents per million input tokens I would not have believed you. Run GLM-5.3-Flash through your evals before the launch discount ends on the 9th, and plug your volumes into the calculator to see what the Fable 5.1 cache change does for you specifically.
Update, September 5: GPT-6 Astra and two more
OpenAI shipped GPT-6 Astra on September 3. The API price is $10 / $50 per million tokens with cached input at $1.00, which is exactly Fable 5's price list and four times what Fable 5.1 charges for a cache read. Two things to know before you budget. Prompts over 272K tokens move to a long-context tier at $20 / $75, and the optional fast mode doubles everything. Access is also staggered: enterprise customers in OpenAI's trusted-access program got it first, with general API availability following over the next days.
On benchmarks OpenAI leads with ARC-AGI-3 and FrontierMath numbers I can't verify. The independent picture is narrower. Artificial Analysis scores it 61.2, which lands it fifth, below Fable 5.1 (65.7), Opus 5 (63.1) and Fable 5 (62.1) and a hair above GPT-5.6 Sol (60.9). Its GPQA Diamond result of 96.1% is the highest AA has measured for any model. At Sol's promotional $4 / $20 the 5.6 flagship is still the better deal for most OpenAI workloads, at least until Astra's price or score moves.
Two smaller launches from September 2. Gemini 3.8 Flash keeps the $0.75 / $3.75 introductory rate of 3.7 Flash (through December 31, then $1.50 / $7.50) and scores 58.7, up from 56.0. Google says it's a tune of 3.7 Flash aimed at coding agents, and 3.7 stays available. Muse Spark 1.3 from Meta is also a same-price update, $1.25 / $4.25 with cached input at $0.15, and it jumps to 60.8 from 1.2's 56.8. That puts an open-API model at Sol's score for a quarter of Sol's blended price. I'd put it on the same eval list as GLM-5.3-Flash.