DeepSeek
DeepSeek V4 Flash 0731
Official DeepSeek V4 Flash release with enhanced agentic capabilities and integrated DSpark speculative decoding
Model facts
- Context window
- 1M tokens
- Maximum output
- 384K tokens
- Cheapest paid input
- $0.03 per 1M tokens
- Cheapest paid output
- $0.1 per 1M tokens
- Open weights
- yes
- Providers
- 53
Capabilities and modalities
reasoning, tool calling, structured output, temperature control, open weights, text.
Local hardware estimate
At 4K context: Q4_K_M 187.2 GiB, Q8_0 328.8 GiB, F16 594.4 GiB working memory.
- Parameters
- 304.2 billion
- Architecture
- deepseekv4
- Model license
- mit
Planning estimate adapted from the MIT-licensed whichllm estimator using metadata from Hugging Face; not a vendor minimum requirement.
Artificial Analysis benchmarks
Compare this model on the full LLM benchmark leaderboard.
| Configuration | AAI | Coding | Math | Output tokens/s |
|---|---|---|---|---|
| Reasoning, Max Effort | 51.8 | 69.1 | not measured | 129.7 |
Observed price history
- 2026-08-24: $0.035 input / $0.07 output per 1M tokens
- 2026-08-26: $0.06 input / $0.12 output per 1M tokens
- 2026-08-27: $0.05 input / $0.1 output per 1M tokens
- 2026-08-28: $0.06 input / $0.12 output per 1M tokens
- 2026-08-29: $0.045 input / $0.09 output per 1M tokens
- 2026-08-30: $0.076 input / $0.153 output per 1M tokens
- 2026-09-01: $0.03 input / $0.1 output per 1M tokens
API providers
- Alibaba Token Plan — model id deepseek-v4-flash-0731; free input / free output per 1M tokens; provider documentation
- Alibaba Token Plan (China) — model id deepseek-v4-flash-0731; free input / free output per 1M tokens; provider documentation
- Nvidia — model id deepseek-ai/deepseek-v4-flash-0731; free input / free output per 1M tokens; provider documentation
- SCNet Token Plan — model id DeepSeek-V4-Flash-0731; free input / free output per 1M tokens; provider documentation
- Eden AI — model id flexai/deepseek-v4-flash-0731; $0.03 input / $0.1 output per 1M tokens; provider documentation
- Vercel AI Gateway — model id deepseek/deepseek-v4-flash-0731; $0.076 input / $0.153 output per 1M tokens; provider documentation
- OpenRouter — model id deepseek/deepseek-v4-flash-0731; $0.065 input / $0.18 output per 1M tokens; provider documentation
- Ambient — model id deepseek/deepseek-v4-flash-0731; $0.08 input / $0.18 output per 1M tokens; provider documentation
- Deep Infra — model id deepseek-ai/DeepSeek-V4-Flash-0731; $0.08 input / $0.18 output per 1M tokens; provider documentation
- Eden AI — model id deepinfra/deepseek-ai/DeepSeek-V4-Flash-0731; $0.08 input / $0.18 output per 1M tokens; provider documentation
- DigitalOcean — model id deepseek-v4-flash-0731; $0.08 input / $0.252 output per 1M tokens; provider documentation
- Baseten — model id deepseek-ai/DeepSeek-V4-Flash-0731; $0.13 input / $0.26 output per 1M tokens; provider documentation
- Perplexity Agent — model id deepseek/deepseek-v4-flash-0731; $0.13 input / $0.26 output per 1M tokens; provider documentation
- RunInfra — model id deepseek-ai/DeepSeek-V4-Flash-0731; $0.13 input / $0.27 output per 1M tokens; provider documentation
- Cortecs — model id deepseek-v4-flash-0731; $0.13 input / $0.28 output per 1M tokens; provider documentation
- Inceptron — model id deepseek-ai/DeepSeek-V4-Flash-0731; $0.13 input / $0.28 output per 1M tokens; provider documentation
- Weights & Biases — model id deepseek-ai/DeepSeek-V4-Flash-0731; $0.13 input / $0.28 output per 1M tokens; provider documentation
- OpenReason — model id deepseek-ai/deepseek-v4-flash-0731; $0.137 input / $0.274 output per 1M tokens; provider documentation
- AMD — model id DeepSeek-V4-Flash; $0.14 input / $0.28 output per 1M tokens; provider documentation
- Eden AI — model id databricks/databricks-deepseek-v4-flash-0731; $0.14 input / $0.28 output per 1M tokens; provider documentation
- Eden AI — model id nebius/deepseek-ai/DeepSeek-V4-Flash-0731; $0.14 input / $0.28 output per 1M tokens; provider documentation
- Eden AI — model id together_ai/deepseek-ai/DeepSeek-V4-Flash-0731; $0.14 input / $0.28 output per 1M tokens; provider documentation
- Hugging Face — model id deepseek-ai/DeepSeek-V4-Flash-0731; $0.14 input / $0.28 output per 1M tokens; provider documentation
- NanoGPT — model id deepseek/deepseek-v4-flash-0731; $0.14 input / $0.28 output per 1M tokens; provider documentation
- Nebius Token Factory — model id deepseek-ai/DeepSeek-V4-Flash-0731; $0.14 input / $0.28 output per 1M tokens; provider documentation
- OpenRouter (batch) — model id deepseek/deepseek-v4-flash-0731:batch; $0.14 input / $0.28 output per 1M tokens; provider documentation
- Requesty — model id deepseek-v4-flash-0731; $0.14 input / $0.28 output per 1M tokens; provider documentation
- Requesty (EU) — model id deepseek-v4-flash-0731@eu; $0.14 input / $0.28 output per 1M tokens; provider documentation
- Together AI — model id deepseek-ai/DeepSeek-V4-Flash-0731; $0.14 input / $0.28 output per 1M tokens; provider documentation
- AIHubMix — model id deepseek-v4-flash-0731; $0.142 input / $0.284 output per 1M tokens; provider documentation
- OrcaRouter — model id deepseek/deepseek-v4-flash-0731; $0.147 input / $0.295 output per 1M tokens; provider documentation
- Venice AI — model id deepseek-v4-flash-0731; $0.175 input / $0.35 output per 1M tokens; provider documentation
- Eden AI — model id tensorx/deepseek/deepseek-v4-flash-0731; $0.25 input / $0.3 output per 1M tokens; provider documentation
- TensorX — model id deepseek/deepseek-v4-flash-0731; $0.25 input / $0.3 output per 1M tokens; provider documentation
- GreenPT — model id deepseek-v4-flash-0731; $0.16 input / $0.399 output per 1M tokens; provider documentation
- Alibaba — model id deepseek-v4-flash-0731; $0.2 input / $0.4 output per 1M tokens; provider documentation
- AKI.IO — model id deepseek-v4-flash-0731-284b; $0.2 input / $0.5 output per 1M tokens; provider documentation
- Eden AI — model id qwen/deepseek-v4-flash-0731; $0.176 input / $0.528 output per 1M tokens; provider documentation
- Merge Gateway — model id deepseek/deepseek-v4-flash-0731-fast; $0.28 input / $0.56 output per 1M tokens; provider documentation
- Eden AI — model id fireworks_ai/accounts/fireworks/models/deepseek-v4-flash-0731; $0.22 input / $0.66 output per 1M tokens; provider documentation
- Fireworks AI — model id accounts/fireworks/models/deepseek-v4-flash-0731; $0.22 input / $0.66 output per 1M tokens; provider documentation
- Merge Gateway — model id deepseek/deepseek-v4-flash-0731; $0.22 input / $0.66 output per 1M tokens; provider documentation
- Tinfoil — model id deepseek-v4-flash; $0.3 input / $0.7 output per 1M tokens; provider documentation
- Eden AI — model id scaleway/deepseek-v4-flash-0731; $0.465 input / $0.929 output per 1M tokens; provider documentation
- Scaleway — model id deepseek-v4-flash-0731; $0.468 input / $0.936 output per 1M tokens; provider documentation
- EmpirioLabs AI — model id deepseek-v4-flash-0731; $0.424 input / $1.27 output per 1M tokens; provider documentation
- Charm Hyper — model id deepseek-v4-flash-0731; $0.44 input / $1.32 output per 1M tokens; provider documentation
- Cloudflare Workers AI — model id @cf/deepseek-ai/deepseek-v4-flash-0731; $0.44 input / $1.32 output per 1M tokens; provider documentation
- Eden AI — model id cloudflare/@cf/deepseek-ai/deepseek-v4-flash-0731; $0.44 input / $1.32 output per 1M tokens; provider documentation
- Kilo Gateway — model id deepseek/deepseek-v4-flash-0731; $0.44 input / $1.32 output per 1M tokens; provider documentation
- Ofox — model id deepseek/deepseek-v4-flash-0731; $0.44 input / $1.32 output per 1M tokens; provider documentation
- Volcengine Ark — model id deepseek-v4-flash-ga-260731; $0.445 input / $1.34 output per 1M tokens; provider documentation
- Ollama Cloud — model id deepseek-v4-flash:0731; not published input / not published output per 1M tokens; provider documentation