DeepSeek V4 Flash — API Pricing & Benchmarks
DeepSeek V4 Flash is DeepSeek's mid reasoning model, released in 2026-07. The API costs $0.44 per 1M input tokens and $1.32 per 1M output tokens ($0.66 blended at a 3:1 ratio), with cached input at $0.01. It supports a 1M-token context window. On the independently measured Artificial Analysis Intelligence Index it scores 51.8, ranking #28 of 68 models tracked on this site, with a median output speed of 136 tokens/sec. Pricing was changed on 2026-08-16 (previously $0.14/$0.28 per 1M tokens).
Explore provider routes and additional benchmark sources for DeepSeek V4 Flash
| Provider | DeepSeek |
| Input price / 1M tokens | $0.44 |
| Output price / 1M tokens | $1.32 |
| Cached input / 1M tokens | $0.01 |
| Blended price (3:1) | $0.66 |
| Context window | 1M |
| Max output tokens | 384K |
| Knowledge cutoff | — |
| License | open weights |
| Input modalities | T |
| Output speed (tok/s) | 136 |
| Released | 2026-07 |
| Artificial Analysis Intelligence Index | 51.8 |
| LMArena Elo | 1432 |
| GPQA Diamond | 90.8% |
| Humanity's Last Exam | 38.6% |
Where DeepSeek V4 Flash fits
On a blended 3:1 basis DeepSeek V4 Flash costs $0.66 per 1M tokens, cheaper than 75% of the 80 priced models in this index. It sits in the mid-tier bracket where most production chat, RAG and summarization workloads run: meaningfully cheaper than flagships while keeping most of their capability. The weights are open, so it can be self-hosted for data-residency or cost reasons instead of consumed via API. Cached input is priced at $0.01 (97% below standard input), which rewards prompt structures with long stable prefixes — see our caching guide.
Price history
| Period | Input $/1M | Output $/1M | Cached $/1M |
|---|---|---|---|
| 2026-07-18 → 2026-08-16 | $0.14 | $0.28 | $0.0028 |
| 2026-08-16 → current | $0.44 | $1.32 | $0.01 |
List-price changes recorded by this site. The full dataset is public: price-history.json (CC BY 4.0).
Alternatives at a similar price
- GPT-4.1 mini (OpenAI) — $0.70/1M blended, AA Index 14.8
- Qwen3.7-Plus (Alibaba Qwen) — $0.70/1M blended, AA Index 39.4
- Mistral Large 3 (Mistral AI) — $0.75/1M blended, AA Index 15.9
- Gemini 3.1 Flash-Lite (Google) — $0.56/1M blended, AA Index 25.6
About DeepSeek
DeepSeek publishes open-weights models at disruptive prices. Even after its August 2026 switch to peak/off-peak billing, the V4 line costs a fraction of Western equivalents, and off-peak hours (everything outside 01:00–04:00 and 06:00–10:00 UTC on weekdays) halve the bill again.