AI API Rankings
Benchmark rankings, prices, context windows and speed for every major AI API — on one page.
Price vs performance
X-axis: blended price (3:1 input:output weighting, log scale). Y-axis: pick any benchmark below. Upper-left is cheap and strong.
Data: Artificial Analysis · official provider pricing pages
Rankings by benchmark
Pick a benchmark to see the top models. Sources are listed in the footer.
Data: Artificial Analysis · LMArena leaderboard dataset (CC BY 4.0) · provider model cards
Best model by use case
Top five models for common workloads, taken from the table below. * marks provider-reported scores.
All models
Click a column header to sort. Prices are official API list prices in USD per 1M tokens.
| Model | Input $/1M | Output $/1M | Blended | Cached input | Context | AA Index | Arena Elo | GPQA | SWE-bench | Speed t/s | Released |
|---|---|---|---|---|---|---|---|---|---|---|---|
| Claude Fable 5.1 Anthropic | $10.0 | $50.0 | $20.0 | $0.25 | 1M | 65.7 | — | 93.7% | — | 71 | 2026-09 |
| Claude Opus 5 Anthropic | $5.00 | $25.0 | $10.0 | $0.50 | 1M | 63.1 | 1505 | 93.2% | 97% | 52 | 2026-07 |
| Claude Fable 5 Anthropic | $10.0 | $50.0 | $20.0 | $1.00 | 1M | 62.1 | 1494 | 92.6% | 95% | 68 | 2026-06 |
| GPT-6 Astra OpenAI | $10.0 | $50.0 | $20.0 | $1.00 | 1.1M | 61.2 | — | 96.1% | — | — | 2026-09 |
| GPT-5.6 Sol OpenAI | $4.00 | $20.0 | $8.00 | $0.40 | 1.1M | 60.9 | 1455 | 94.1% | — | 81 | 2026-07 |
| Grok 4.6 xAI | $2.00 | $6.00 | $3.00 | $0.50 | 500K | 60.9 | 1444 | 94.9% | — | 60 | 2026-08 |
| Muse Spark 1.3 Meta | $1.25 | $4.25 | $2.00 | $0.15 | 1.0M | 60.8 | — | 94.1% | — | 115 | 2026-09 |
| Kimi K3 Moonshot AI | $3.00 | $15.0 | $6.00 | $0.30 | 1.0M | 59.7 | 1476 | 93.5% | 93.4% | 38 | 2026-07 |
| GLM-5.3 Zhipu AI (Z.ai) | $1.40 | $4.40 | $2.15 | $0.26 | 1M | 59.5 | 1474 | 91.7% | — | 78 | 2026-08 |
| Gemini 3.8 Flash Google | $0.75 | $3.75 | $1.50 | $0.07 | 1.0M | 58.7 | — | 95.3% | — | 417 | 2026-09 |
| Qwen3.8-Max Alibaba Qwen | $2.00 | $6.00 | $3.00 | $0.25 | 1M | 58.1 | 1480 | 92.7% | — | 39 | 2026-08 |
| GLM-5.3-Flash Zhipu AI (Z.ai) | $0.15 | $0.50 | $0.24 | $0.03 | 1M | 57.5 | 1471 | 91.2% | — | 45 | 2026-08 |
| Claude Opus 4.8 Anthropic | $5.00 | $25.0 | $10.0 | $0.50 | 1M | 57.3 | 1461 | 92% | 88.6%* | 63 | 2026-05 |
| Muse Spark 1.2 Meta | $1.25 | $4.25 | $2.00 | $0.15 | 1.0M | 56.8 | 1488 | 90.4% | — | 224 | 2026-08 |
| GPT-5.6 Terra OpenAI | $2.00 | $12.0 | $4.50 | $0.20 | 1.1M | 56.6 | 1447 | 92.5% | — | 104 | 2026-07 |
| GPT-5.5 OpenAI | $5.00 | $30.0 | $11.3 | $0.50 | 1.1M | 56.3 | 1472 | 93.5% | 88.7%* | 66 | 2026-04 |
| Gemini 3.7 Flash Google | $0.75 | $3.75 | $1.50 | $0.07 | 1.0M | 56 | 1491 | 94.5% | — | 309 | 2026-08 |
| Grok 4.5 xAI | $2.00 | $6.00 | $3.00 | $0.30 | 500K | 55.8 | 1452 | 93.1% | 86.6% | 53 | 2026-07 |
| Qwen3.8-Flash Alibaba Qwen | $0.15 | $0.47 | $0.23 | $0.01 | 1M | 55.8 | — | 92.3% | — | 86 | 2026-08 |
| Claude Sonnet 5 Anthropic | $2.00 | $10.0 | $4.00 | $0.20 | 1M | 55.3 | 1443 | 91.1% | 85.2% | 81 | 2026-06 |
| Claude Opus 4.7 Anthropic | $5.00 | $25.0 | $10.0 | $0.50 | 1M | 55 | 1490 | 91.4% | 87.6%* | 46 | 2026-04 |
| DeepSeek V4 Pro DeepSeek | $1.32 | $3.96 | $1.98 | $0.04 | 1M | 53.2 | 1440 | 92.8% | 96.4% | 67 | 2026-08 |
| Muse Spark 1.1 Meta | $1.25 | $4.25 | $2.00 | — | — | 53.2 | 1479 | 89.8% | — | 189 | 2026-07 |
| GPT-5.4 OpenAI | $2.50 | $15.0 | $5.63 | $0.25 | 1.1M | 53.1 | 1470 | 92% | 76.9% | 150 | 2026-03 |
| GLM-5.2 Zhipu AI (Z.ai) | $1.40 | $4.40 | $2.15 | $0.26 | 1M | 52.6 | 1467 | 89.5% | — | 71 | 2026-06 |
| GPT-5.6 Luna OpenAI | $0.20 | $1.20 | $0.45 | $0.02 | 1.1M | 52.3 | 1430 | 91.1% | — | 110 | 2026-07 |
| Gemini 3.5 Flash Google | $1.50 | $9.00 | $3.38 | $0.15 | 1.0M | 52 | 1484 | 92.2% | 79.3% | 251 | 2026-05 |
| DeepSeek V4 Flash DeepSeek | $0.44 | $1.32 | $0.66 | $0.01 | 1M | 51.8 | 1432 | 90.8% | — | 136 | 2026-07 |
| Gemini 3.6 Flash Google | $0.75 | $3.75 | $1.50 | $0.07 | 1.0M | 51.6 | 1476 | 92.8% | — | 196 | 2026-07 |
| Claude Sonnet 4.6 Anthropic | $3.00 | $15.0 | $6.00 | $0.30 | 1M | 48.4 | 1458 | 87.5% | 79.6% | 49 | 2026-02 |
| Gemini 3.1 Pro Google | $2.00 | $12.0 | $4.50 | $0.20 | 1.0M | 47.7 | 1480 | 94.1% | 80.6% | 118 | 2026-02 |
| Qwen3.7-Max Alibaba Qwen | $2.50 | $7.50 | $3.75 | $0.25 | 1M | 46.7 | 1474 | 92.3% | 80.4% | 196 | 2026-05 |
| GPT-5.3 Codex OpenAI | $1.75 | $14.0 | $4.81 | $0.17 | 400K | 45.5 | — | 91.5% | 85% | 126 | — |
| MiniMax M3 MiniMax | $0.30 | $1.20 | $0.52 | $0.06 | 1M | 45.4 | 1435 | 92.9% | 80.5% | 85 | 2026-06 |
| Kimi K2.6 Moonshot AI | $0.95 | $4.00 | $1.71 | $0.16 | 262K | 45.1 | 1455 | 91.1% | 80.2%* | 41 | 2026-04 |
| Claude Opus 4.6 Anthropic | $5.00 | $25.0 | $10.0 | $0.50 | 1M | 44.9 | 1503 | 89.6% | — | 46 | — |
| Kimi K2.7 Code Moonshot AI | $0.95 | $4.00 | $1.71 | $0.19 | 262K | 43 | — | 89.6% | — | 39 | 2026-06 |
| Claude Opus 4.5 Anthropic | $5.00 | $25.0 | $10.0 | $0.50 | 200K | 41.9 | 1450 | 86.6% | — | — | 2025-11 |
| Solar Pro 4 Upstage | $0.30 | $1.20 | $0.52 | $0.06 | 524K | 41.6 | 1377 | 89.1% | — | 62 | 2026-08 |
| GPT-5.4 mini OpenAI | $0.75 | $4.50 | $1.69 | $0.07 | 400K | 40.9 | 1412 | 87.5% | — | 175 | 2026-03 |
| GLM-5 Zhipu AI (Z.ai) | $1.00 | $3.20 | $1.55 | $0.20 | 200K | 40.6 | 1446 | 82% | — | 48 | 2026-02 |
| GPT-5.4 nano OpenAI | $0.20 | $1.25 | $0.46 | $0.02 | 400K | 39.7 | 1373 | 81.7% | — | 163 | 2026-03 |
| Qwen3.7-Plus Alibaba Qwen | $0.40 | $1.60 | $0.70 | — | 1M | 39.4 | 1454 | 90% | 77.7% | 56 | 2026-06 |
| MiniMax M2.7 MiniMax | $0.30 | $1.20 | $0.52 | $0.06 | — | 38.9 | 1405 | 87.4% | — | 50 | — |
| Grok 4.3 xAI | $1.25 | $2.50 | $1.56 | $0.20 | 1M | 37.9 | 1398 | 90.1% | — | 137 | 2026-04 |
| Gemini 3.5 Flash-Lite Google | $0.30 | $2.50 | $0.85 | $0.03 | 1.0M | 37.4 | 1436 | 83.8% | — | 366 | 2026-07 |
| GLM-4.7 Zhipu AI (Z.ai) | $0.60 | $2.20 | $1.00 | $0.11 | — | 34.5 | 1436 | 85.9% | — | 113 | — |
| Mistral Medium 3.5 Mistral AI | $1.50 | $7.50 | $3.00 | — | — | 30.4 | 1421 | 74.8% | 77.6% | 131 | 2026-04 |
| Gemini 2.5 Pro Google | $1.25 | $10.0 | $3.44 | $0.13 | 1.0M | 25.9 | 1458 | 84.4% | — | 124 | 2025-06 |
| Gemini 3.1 Flash-Lite Google | $0.25 | $1.50 | $0.56 | $0.03 | 1.0M | 25.6 | 1415 | 82.2% | — | 276 | 2026-03 |
| Gemini 2.5 Flash Google | $0.30 | $2.50 | $0.85 | $0.03 | 1.0M | 24.2 | 1417 | 79.3% | — | 215 | 2025-06 |
| Command A+ Cohere | — | — | — | — | 128K | 22.8 | — | 76.1% | — | 262 | 2026-05 |
| ERNIE 5.0 Baidu | $0.60 | $2.10 | $0.97 | — | 128K | 22.3 | 1444 | 77.7% | — | — | — |
| Nova 2 Pro (preview) Amazon | $1.25 | $10.0 | $3.44 | $0.31 | 1M | 22.1 | — | 78.5% | — | 128 | 2025-12 |
| Nova 2 Omni (preview) Amazon | $0.30 | $2.50 | $0.85 | — | 1M | 21.3 | — | 76%* | — | — | 2025-12 |
| Nova 2 Lite Amazon | $0.30 | $2.50 | $0.85 | $0.07 | 1M | 20.8 | 1363 | 81.1% | — | 191 | 2025-12 |
| Mistral Small 4 Mistral AI | $0.15 | $0.60 | $0.26 | — | — | 19.7 | — | 76.9% | — | 164 | 2026-03 |
| GPT-4.1 OpenAI | $2.00 | $8.00 | $3.50 | $0.50 | 1.0M | 19.6 | 1383 | 66.6% | — | — | 2025-04 |
| GLM-4.5-Air Zhipu AI (Z.ai) | $0.20 | $1.10 | $0.43 | $0.03 | — | 16.7 | 1384 | 73.3% | — | 81 | 2025-07 |
| Mistral Large 3 Mistral AI | $0.50 | $1.50 | $0.75 | — | — | 15.9 | 1428 | 68% | — | 74 | 2025-12 |
| GPT-4.1 mini OpenAI | $0.40 | $1.60 | $0.70 | $0.10 | 1.0M | 14.8 | 1340 | 66.4% | — | — | 2025-04 |
| Solar Pro 3 Upstage | $0.15 | $0.60 | $0.26 | $0.01 | 128K | 14.5 | — | 72.4% | — | 142 | 2026-01 |
| Solar Pro 2 Upstage | $0.15 | $0.60 | $0.26 | $0.01 | — | 12.5 | — | 57.8% | — | — | 2025-05 |
| GPT-4o OpenAI | $2.50 | $10.0 | $4.38 | $1.25 | 128K | 12.3 | — | 65.5% | — | — | 2024-05 |
| GPT-4.1 nano OpenAI | $0.10 | $0.40 | $0.18 | $0.03 | 1.0M | 9.6 | — | 51.2% | — | — | 2025-04 |
| Command A Cohere | $2.50 | $10.0 | $4.38 | — | 256K | 7.5 | — | 52.7% | — | 57 | 2025-03 |
| GPT-4o mini OpenAI | $0.15 | $0.60 | $0.26 | $0.07 | 128K | 6.7 | — | 42.6% | — | — | 2024-07 |
| Nova Micro Amazon | $0.04 | $0.14 | $0.06 | — | 128K | 4.4 | — | 35.8% | — | 288 | 2024-12 |
| GPT-5.5 Pro OpenAI | $30.0 | $180.0 | $67.5 | — | 1.1M | — | — | 93.9%* | — | — | 2026-04 |
| GPT-5.4 Pro OpenAI | $30.0 | $180.0 | $67.5 | — | 1.1M | — | — | — | — | — | 2026-03 |
| Claude Haiku 4.5 Anthropic | $1.00 | $5.00 | $2.00 | $0.10 | 200K | — | 1395 | — | 73.3% | — | 2025-10 |
| Claude Sonnet 4.5 Anthropic | $3.00 | $15.0 | $6.00 | $0.30 | 200K | — | 1438 | — | — | — | 2025-09 |
| Grok Build 0.1 xAI | $1.00 | $2.00 | $1.25 | $0.20 | 256K | — | — | — | — | — | — |
| Grok 4.20 xAI | $1.25 | $2.50 | $1.56 | $0.20 | 1M | — | 1451 | — | 76.7%* | — | 2026-03 |
| Qwen3.6-Flash Alibaba Qwen | $0.25 | $1.50 | $0.56 | — | 1M | — | — | — | — | — | — |
| Magistral Medium Mistral AI | $2.00 | $5.00 | $2.75 | — | — | — | — | — | — | — | — |
| Codestral 25.08 Mistral AI | $0.30 | $0.90 | $0.45 | — | — | — | — | — | — | — | 2025-08 |
| ERNIE 5.1 Baidu | $0.59 | $2.65 | $1.10 | — | 128K | — | 1468 | — | — | — | 2026-05 |
| Command R7B Cohere | $0.04 | $0.15 | $0.07 | — | 128K | — | — | — | — | — | 2024-12 |
| HyperCLOVA X HCX-007 Naver | $0.84 | $3.37 | $1.47 | — | 128K | — | — | — | — | — | — |
| K-EXAONE 236B LG AI Research | $0.20 | $0.80 | $0.35 | $0.10 | 262K | — | — | — | — | — | — |
Reading the table (promo prices, self-reported scores, etc.)
Compare models head-to-head
Pick up to 4 models to compare prices, specs and benchmark scores side by side.
Data: Artificial Analysis · LMArena dataset (CC BY 4.0) · provider model cards & pricing pages
Frontier progress over time
Intelligence Index by release month. Each dot is a model, colored by provider.
Data: Artificial Analysis
Monthly cost calculator
Enter your expected usage to estimate monthly cost per model, cheapest first.
Estimate from a prompt…
Chat apps typically see 3–5× more input than output. Cache hit rate applies the provider's cached-input price to that share of input; models without cache pricing use the full input price.
Benchmark guide
What each benchmark used on this page actually measures.
AA Intelligence Index Composite score
Artificial Analysis independently re-runs a battery of evaluations (knowledge, math, coding, instruction following) against each model's API and combines them into one number. Our default ranking metric; gaps under ~2 points are noise.
LMArena Elo Human preference
Chess-style rating from blind A/B votes on real user prompts. Measures perceived answer quality — helpfulness, tone, formatting — rather than raw reasoning ability.
GPQA Diamond Graduate science
Google-proof graduate-level science questions. Frontier models now exceed 90%, so it separates mid-tier from frontier but no longer separates frontier models from each other.
AIME Competition math
American Invitational Mathematics Examination problems with integer answers. Top models score 98–100%, so treat it as a pass/fail signal rather than a ranking.
SWE-bench Verified Agentic coding
Real GitHub issues: the model must produce a patch that passes the repository's tests. The closest benchmark to real software work; sensitive to the agent harness used.
LiveCodeBench Algorithmic coding
Competitive-programming problems collected continuously after training cutoffs to avoid contamination. Measures raw problem-solving, complementing SWE-bench.
MMLU-Pro Broad knowledge
A harder, 10-choice successor to MMLU covering 14 subject areas. Less saturated than the original, still partly a memorization test.
Humanity's Last Exam Frontier exam
~2,500 expert-written questions designed to stay hard as older benchmarks saturate. Frontier scores of 40–58% make it one of the most discriminating tests available.
MMMU Multimodal
College-level questions requiring charts, diagrams and images to answer. The main signal for vision-input quality.
Frequently asked questions
Where does the pricing data come from?
Prices come from each provider's official pricing page (OpenAI Platform, Anthropic Docs, Google AI for Developers, xAI Docs, DeepSeek API Docs, and others). Listed prices are standard-tier list prices per 1M tokens; cached-input, batch and volume discounts are not applied unless noted.
Are the benchmark scores official?
Where possible we prefer independently re-measured numbers such as Artificial Analysis; otherwise we use provider-reported scores. The same benchmark can vary with run settings (tools, reasoning budget), so treat scores as directional rather than absolute.
What is a blended price?
A weighted average of input and output token prices at a 3:1 input-to-output ratio, a common industry convention for comparing model prices with a single number.
How often is the data updated?
Check the data-as-of date at the top and bottom of the page. New models and price changes are verified manually before being added. If you spot an error, please report it.
Which model should I choose?
As a rule of thumb: frontier models for complex reasoning and coding (expensive), small models for high-volume simple tasks (cheap), and mid-tier models in between. Models in the upper-left of the price-vs-performance chart offer the best value.