Blog
Hand-written guides to LLM API pricing and benchmarks, plus a monthly log of price changes across every provider we track.
LLM API Price Changes — September 2026
Claude Fable 5.1 cut cache reads by 75%, OpenAI trimmed GPT-5.6 Sol, DeepSeek moved to peak-hour billing, and Sonnet 5's launch price became permanent. Every verified change since the August log.
LLM API Price Changes — August 2026
OpenAI cut GPT-5.6 Luna by 80% and Terra by 20%, and Qwen3.8-Max entered the top five. This month's verified price changes and what they mean if you're buying.
How LLM API Pricing Actually Works: Tokens, Caching, Batch and Blended Prices
Input vs output tokens, prompt caching, batch discounts, reasoning tokens and blended prices. A practical guide to reading an LLM pricing page and predicting your real bill.
GPQA, AIME, SWE-bench, Elo: What LLM Benchmarks Actually Measure
A plain-language guide to the benchmarks used to rank LLMs: what each one tests, what a good score means, and the caveats (contamination, self-reporting, saturation) worth knowing.
How to Choose an LLM API in 2026: A Practical Decision Guide
A decision framework for picking an LLM API: match the model tier to the job, apply hard constraints, find the value frontier, then eval a shortlist on your own tasks.
Context Windows and Prompt Caching: The Two Numbers That Decide Your LLM Bill
Long context is re-billed on every request, and prompt caching discounts repeated tokens by 90 to 99%. How to structure prompts and sessions around both.