AI API Rankings

Benchmark rankings, prices, context windows and speed for every major AI API — on one page.

81 models·17 providers·9 benchmarks·data as of 2026-09-05
Explore more provider routes and benchmark results with source dates and evaluation settings. Public data explorer

Price vs performance

X-axis: blended price (3:1 input:output weighting, log scale). Y-axis: pick any benchmark below. Upper-left is cheap and strong.

Y-axis benchmark

Data: Artificial Analysis · official provider pricing pages

Rankings by benchmark

Pick a benchmark to see the top models. Sources are listed in the footer.

Data: Artificial Analysis · LMArena leaderboard dataset (CC BY 4.0) · provider model cards

Best model by use case

Top five models for common workloads, taken from the table below. * marks provider-reported scores.

All models

Click a column header to sort. Prices are official API list prices in USD per 1M tokens.

ModelInput $/1MOutput $/1MBlendedCached inputContextAA IndexArena EloGPQASWE-benchSpeed t/sReleased
Claude Fable 5.1 Anthropic$10.0$50.0$20.0$0.251M65.793.7%712026-09
Claude Opus 5 Anthropic$5.00$25.0$10.0$0.501M63.1150593.2%97%522026-07
Claude Fable 5 Anthropic$10.0$50.0$20.0$1.001M62.1149492.6%95%682026-06
GPT-6 Astra OpenAI$10.0$50.0$20.0$1.001.1M61.296.1%2026-09
GPT-5.6 Sol OpenAI$4.00$20.0$8.00$0.401.1M60.9145594.1%812026-07
Grok 4.6 xAI$2.00$6.00$3.00$0.50500K60.9144494.9%602026-08
Muse Spark 1.3 Meta$1.25$4.25$2.00$0.151.0M60.894.1%1152026-09
Kimi K3 Moonshot AI$3.00$15.0$6.00$0.301.0M59.7147693.5%93.4%382026-07
GLM-5.3 Zhipu AI (Z.ai)$1.40$4.40$2.15$0.261M59.5147491.7%782026-08
Gemini 3.8 Flash Google$0.75$3.75$1.50$0.071.0M58.795.3%4172026-09
Qwen3.8-Max Alibaba Qwen$2.00$6.00$3.00$0.251M58.1148092.7%392026-08
GLM-5.3-Flash Zhipu AI (Z.ai)$0.15$0.50$0.24$0.031M57.5147191.2%452026-08
Claude Opus 4.8 Anthropic$5.00$25.0$10.0$0.501M57.3146192%88.6%*632026-05
Muse Spark 1.2 Meta$1.25$4.25$2.00$0.151.0M56.8148890.4%2242026-08
GPT-5.6 Terra OpenAI$2.00$12.0$4.50$0.201.1M56.6144792.5%1042026-07
GPT-5.5 OpenAI$5.00$30.0$11.3$0.501.1M56.3147293.5%88.7%*662026-04
Gemini 3.7 Flash Google$0.75$3.75$1.50$0.071.0M56149194.5%3092026-08
Grok 4.5 xAI$2.00$6.00$3.00$0.30500K55.8145293.1%86.6%532026-07
Qwen3.8-Flash Alibaba Qwen$0.15$0.47$0.23$0.011M55.892.3%862026-08
Claude Sonnet 5 Anthropic$2.00$10.0$4.00$0.201M55.3144391.1%85.2%812026-06
Claude Opus 4.7 Anthropic$5.00$25.0$10.0$0.501M55149091.4%87.6%*462026-04
DeepSeek V4 Pro DeepSeek$1.32$3.96$1.98$0.041M53.2144092.8%96.4%672026-08
Muse Spark 1.1 Meta$1.25$4.25$2.0053.2147989.8%1892026-07
GPT-5.4 OpenAI$2.50$15.0$5.63$0.251.1M53.1147092%76.9%1502026-03
GLM-5.2 Zhipu AI (Z.ai)$1.40$4.40$2.15$0.261M52.6146789.5%712026-06
GPT-5.6 Luna OpenAI$0.20$1.20$0.45$0.021.1M52.3143091.1%1102026-07
Gemini 3.5 Flash Google$1.50$9.00$3.38$0.151.0M52148492.2%79.3%2512026-05
DeepSeek V4 Flash DeepSeek$0.44$1.32$0.66$0.011M51.8143290.8%1362026-07
Gemini 3.6 Flash Google$0.75$3.75$1.50$0.071.0M51.6147692.8%1962026-07
Claude Sonnet 4.6 Anthropic$3.00$15.0$6.00$0.301M48.4145887.5%79.6%492026-02
Gemini 3.1 Pro Google$2.00$12.0$4.50$0.201.0M47.7148094.1%80.6%1182026-02
Qwen3.7-Max Alibaba Qwen$2.50$7.50$3.75$0.251M46.7147492.3%80.4%1962026-05
GPT-5.3 Codex OpenAI$1.75$14.0$4.81$0.17400K45.591.5%85%126
MiniMax M3 MiniMax$0.30$1.20$0.52$0.061M45.4143592.9%80.5%852026-06
Kimi K2.6 Moonshot AI$0.95$4.00$1.71$0.16262K45.1145591.1%80.2%*412026-04
Claude Opus 4.6 Anthropic$5.00$25.0$10.0$0.501M44.9150389.6%46
Kimi K2.7 Code Moonshot AI$0.95$4.00$1.71$0.19262K4389.6%392026-06
Claude Opus 4.5 Anthropic$5.00$25.0$10.0$0.50200K41.9145086.6%2025-11
Solar Pro 4 Upstage$0.30$1.20$0.52$0.06524K41.6137789.1%622026-08
GPT-5.4 mini OpenAI$0.75$4.50$1.69$0.07400K40.9141287.5%1752026-03
GLM-5 Zhipu AI (Z.ai)$1.00$3.20$1.55$0.20200K40.6144682%482026-02
GPT-5.4 nano OpenAI$0.20$1.25$0.46$0.02400K39.7137381.7%1632026-03
Qwen3.7-Plus Alibaba Qwen$0.40$1.60$0.701M39.4145490%77.7%562026-06
MiniMax M2.7 MiniMax$0.30$1.20$0.52$0.0638.9140587.4%50
Grok 4.3 xAI$1.25$2.50$1.56$0.201M37.9139890.1%1372026-04
Gemini 3.5 Flash-Lite Google$0.30$2.50$0.85$0.031.0M37.4143683.8%3662026-07
GLM-4.7 Zhipu AI (Z.ai)$0.60$2.20$1.00$0.1134.5143685.9%113
Mistral Medium 3.5 Mistral AI$1.50$7.50$3.0030.4142174.8%77.6%1312026-04
Gemini 2.5 Pro Google$1.25$10.0$3.44$0.131.0M25.9145884.4%1242025-06
Gemini 3.1 Flash-Lite Google$0.25$1.50$0.56$0.031.0M25.6141582.2%2762026-03
Gemini 2.5 Flash Google$0.30$2.50$0.85$0.031.0M24.2141779.3%2152025-06
Command A+ Cohere128K22.876.1%2622026-05
ERNIE 5.0 Baidu$0.60$2.10$0.97128K22.3144477.7%
Nova 2 Pro (preview) Amazon$1.25$10.0$3.44$0.311M22.178.5%1282025-12
Nova 2 Omni (preview) Amazon$0.30$2.50$0.851M21.376%*2025-12
Nova 2 Lite Amazon$0.30$2.50$0.85$0.071M20.8136381.1%1912025-12
Mistral Small 4 Mistral AI$0.15$0.60$0.2619.776.9%1642026-03
GPT-4.1 OpenAI$2.00$8.00$3.50$0.501.0M19.6138366.6%2025-04
GLM-4.5-Air Zhipu AI (Z.ai)$0.20$1.10$0.43$0.0316.7138473.3%812025-07
Mistral Large 3 Mistral AI$0.50$1.50$0.7515.9142868%742025-12
GPT-4.1 mini OpenAI$0.40$1.60$0.70$0.101.0M14.8134066.4%2025-04
Solar Pro 3 Upstage$0.15$0.60$0.26$0.01128K14.572.4%1422026-01
Solar Pro 2 Upstage$0.15$0.60$0.26$0.0112.557.8%2025-05
GPT-4o OpenAI$2.50$10.0$4.38$1.25128K12.365.5%2024-05
GPT-4.1 nano OpenAI$0.10$0.40$0.18$0.031.0M9.651.2%2025-04
Command A Cohere$2.50$10.0$4.38256K7.552.7%572025-03
GPT-4o mini OpenAI$0.15$0.60$0.26$0.07128K6.742.6%2024-07
Nova Micro Amazon$0.04$0.14$0.06128K4.435.8%2882024-12
GPT-5.5 Pro OpenAI$30.0$180.0$67.51.1M93.9%*2026-04
GPT-5.4 Pro OpenAI$30.0$180.0$67.51.1M2026-03
Claude Haiku 4.5 Anthropic$1.00$5.00$2.00$0.10200K139573.3%2025-10
Claude Sonnet 4.5 Anthropic$3.00$15.0$6.00$0.30200K14382025-09
Grok Build 0.1 xAI$1.00$2.00$1.25$0.20256K
Grok 4.20 xAI$1.25$2.50$1.56$0.201M145176.7%*2026-03
Qwen3.6-Flash Alibaba Qwen$0.25$1.50$0.561M
Magistral Medium Mistral AI$2.00$5.00$2.75
Codestral 25.08 Mistral AI$0.30$0.90$0.452025-08
ERNIE 5.1 Baidu$0.59$2.65$1.10128K14682026-05
Command R7B Cohere$0.04$0.15$0.07128K2024-12
HyperCLOVA X HCX-007 Naver$0.84$3.37$1.47128K
K-EXAONE 236B LG AI Research$0.20$0.80$0.35$0.10262K

Reading the table (promo prices, self-reported scores, etc.)

    Compare models head-to-head

    Pick up to 4 models to compare prices, specs and benchmark scores side by side.

    Data: Artificial Analysis · LMArena dataset (CC BY 4.0) · provider model cards & pricing pages

    Frontier progress over time

    Intelligence Index by release month. Each dot is a model, colored by provider.

    Monthly cost calculator

    Enter your expected usage to estimate monthly cost per model, cheapest first.

    Estimate from a prompt…

    Chat apps typically see 3–5× more input than output. Cache hit rate applies the provider's cached-input price to that share of input; models without cache pricing use the full input price.

    Benchmark guide

    What each benchmark used on this page actually measures.

    AA Intelligence Index Composite score

    Artificial Analysis independently re-runs a battery of evaluations (knowledge, math, coding, instruction following) against each model's API and combines them into one number. Our default ranking metric; gaps under ~2 points are noise.

    LMArena Elo Human preference

    Chess-style rating from blind A/B votes on real user prompts. Measures perceived answer quality — helpfulness, tone, formatting — rather than raw reasoning ability.

    GPQA Diamond Graduate science

    Google-proof graduate-level science questions. Frontier models now exceed 90%, so it separates mid-tier from frontier but no longer separates frontier models from each other.

    AIME Competition math

    American Invitational Mathematics Examination problems with integer answers. Top models score 98–100%, so treat it as a pass/fail signal rather than a ranking.

    SWE-bench Verified Agentic coding

    Real GitHub issues: the model must produce a patch that passes the repository's tests. The closest benchmark to real software work; sensitive to the agent harness used.

    LiveCodeBench Algorithmic coding

    Competitive-programming problems collected continuously after training cutoffs to avoid contamination. Measures raw problem-solving, complementing SWE-bench.

    MMLU-Pro Broad knowledge

    A harder, 10-choice successor to MMLU covering 14 subject areas. Less saturated than the original, still partly a memorization test.

    Humanity's Last Exam Frontier exam

    ~2,500 expert-written questions designed to stay hard as older benchmarks saturate. Frontier scores of 40–58% make it one of the most discriminating tests available.

    MMMU Multimodal

    College-level questions requiring charts, diagrams and images to answer. The main signal for vision-input quality.

    Frequently asked questions

    Where does the pricing data come from?

    Prices come from each provider's official pricing page (OpenAI Platform, Anthropic Docs, Google AI for Developers, xAI Docs, DeepSeek API Docs, and others). Listed prices are standard-tier list prices per 1M tokens; cached-input, batch and volume discounts are not applied unless noted.

    Are the benchmark scores official?

    Where possible we prefer independently re-measured numbers such as Artificial Analysis; otherwise we use provider-reported scores. The same benchmark can vary with run settings (tools, reasoning budget), so treat scores as directional rather than absolute.

    What is a blended price?

    A weighted average of input and output token prices at a 3:1 input-to-output ratio, a common industry convention for comparing model prices with a single number.

    How often is the data updated?

    Check the data-as-of date at the top and bottom of the page. New models and price changes are verified manually before being added. If you spot an error, please report it.

    Which model should I choose?

    As a rule of thumb: frontier models for complex reasoning and coding (expensive), small models for high-volume simple tasks (cheap), and mid-tier models in between. Models in the upper-left of the price-vs-performance chart offer the best value.