Gemini 3.1 Pro — API Pricing & Benchmarks
Gemini 3.1 Pro is Google's flagship reasoning model, released in 2026-02. The API costs $2.00 per 1M input tokens and $12.0 per 1M output tokens ($4.50 blended at a 3:1 ratio), with cached input at $0.20. It supports a 1.0M-token context window. On the independently measured Artificial Analysis Intelligence Index it scores 47.7, ranking #31 of 68 models tracked on this site, with a median output speed of 118 tokens/sec.
Explore provider routes and additional benchmark sources for Gemini 3.1 Pro
| Provider | |
| Input price / 1M tokens | $2.00 |
| Output price / 1M tokens | $12.0 |
| Cached input / 1M tokens | $0.20 |
| Blended price (3:1) | $4.50 |
| Context window | 1.0M |
| Max output tokens | 66K |
| Knowledge cutoff | 2025-01 |
| License | proprietary |
| Input modalities | T+I+A+V |
| Output speed (tok/s) | 118 |
| Released | 2026-02 |
| Artificial Analysis Intelligence Index | 47.7 |
| LMArena Elo | 1480 |
| GPQA Diamond | 94.1% |
| SWE-bench Verified | 80.6% |
| MMLU-Pro | 91% |
| Humanity's Last Exam | 47% |
| MMMU | 80.5% |
Where Gemini 3.1 Pro fits
On a blended 3:1 basis Gemini 3.1 Pro costs $4.50 per 1M tokens, cheaper than 21% of the 80 priced models in this index. As Google's flagship it is aimed at the hardest workloads — complex reasoning, agentic coding and multi-step tool use — where output quality dominates cost. Cached input is priced at $0.20 (90% below standard input), which rewards prompt structures with long stable prefixes — see our caching guide.
Price history
| Period | Input $/1M | Output $/1M | Cached $/1M |
|---|---|---|---|
| 2026-07-18 → current | $2.00 | $12.0 | $0.20 |
No list-price changes recorded since tracking began. Full dataset: price-history.json (CC BY 4.0).
Alternatives at a similar price
- GPT-5.6 Terra (OpenAI) — $4.50/1M blended, AA Index 56.6
- GPT-4o (OpenAI) — $4.38/1M blended, AA Index 12.3
- Command A (Cohere) — $4.38/1M blended, AA Index 7.5
- GPT-5.3 Codex (OpenAI) — $4.81/1M blended, AA Index 45.5
About Google
Google's Gemini models are natively multimodal — most accept text, images, audio and video input — with 1M-token context windows across the range. The Flash line now ships on a roughly monthly cadence (3.6, 3.7, 3.8 Flash between July and September 2026) at the same introductory price, competing on price-performance and speed rather than peak capability.