Blog

Hand-written guides to LLM API pricing and benchmarks, plus a monthly log of price changes across every provider we track.

LLM API Price Changes — September 2026

Price watch · September 2, 2026

Claude Fable 5.1 cut cache reads by 75%, OpenAI trimmed GPT-5.6 Sol, DeepSeek moved to peak-hour billing, and Sonnet 5's launch price became permanent. Every verified change since the August log.

LLM API Price Changes — August 2026

Price watch · August 9, 2026

OpenAI cut GPT-5.6 Luna by 80% and Terra by 20%, and Qwen3.8-Max entered the top five. This month's verified price changes and what they mean if you're buying.

How LLM API Pricing Actually Works: Tokens, Caching, Batch and Blended Prices

Guide · August 9, 2026

Input vs output tokens, prompt caching, batch discounts, reasoning tokens and blended prices. A practical guide to reading an LLM pricing page and predicting your real bill.

GPQA, AIME, SWE-bench, Elo: What LLM Benchmarks Actually Measure

Guide · August 9, 2026

A plain-language guide to the benchmarks used to rank LLMs: what each one tests, what a good score means, and the caveats (contamination, self-reporting, saturation) worth knowing.

How to Choose an LLM API in 2026: A Practical Decision Guide

Guide · August 9, 2026

A decision framework for picking an LLM API: match the model tier to the job, apply hard constraints, find the value frontier, then eval a shortlist on your own tasks.

Context Windows and Prompt Caching: The Two Numbers That Decide Your LLM Bill

Guide · August 9, 2026

Long context is re-billed on every request, and prompt caching discounts repeated tokens by 90 to 99%. How to structure prompts and sessions around both.