All AI models · LLM benchmarks · Methodology

DeepSeek

DeepSeek V4 Flash 0731

Official DeepSeek V4 Flash release with enhanced agentic capabilities and integrated DSpark speculative decoding

Model facts

Context window
1M tokens
Maximum output
384K tokens
Cheapest paid input
$0.03 per 1M tokens
Cheapest paid output
$0.1 per 1M tokens
Open weights
yes
Providers
53

Capabilities and modalities

reasoning, tool calling, structured output, temperature control, open weights, text.

Local hardware estimate

At 4K context: Q4_K_M 187.2 GiB, Q8_0 328.8 GiB, F16 594.4 GiB working memory.

Parameters
304.2 billion
Architecture
deepseekv4
Model license
mit

Planning estimate adapted from the MIT-licensed whichllm estimator using metadata from Hugging Face; not a vendor minimum requirement.

Artificial Analysis benchmarks

Compare this model on the full LLM benchmark leaderboard.

ConfigurationAAICodingMathOutput tokens/s
Reasoning, Max Effort51.869.1not measured129.7

Observed price history

  • 2026-08-24: $0.035 input / $0.07 output per 1M tokens
  • 2026-08-26: $0.06 input / $0.12 output per 1M tokens
  • 2026-08-27: $0.05 input / $0.1 output per 1M tokens
  • 2026-08-28: $0.06 input / $0.12 output per 1M tokens
  • 2026-08-29: $0.045 input / $0.09 output per 1M tokens
  • 2026-08-30: $0.076 input / $0.153 output per 1M tokens
  • 2026-09-01: $0.03 input / $0.1 output per 1M tokens

API providers

  • Alibaba Token Plan — model id deepseek-v4-flash-0731; free input / free output per 1M tokens; provider documentation
  • Alibaba Token Plan (China) — model id deepseek-v4-flash-0731; free input / free output per 1M tokens; provider documentation
  • Nvidia — model id deepseek-ai/deepseek-v4-flash-0731; free input / free output per 1M tokens; provider documentation
  • SCNet Token Plan — model id DeepSeek-V4-Flash-0731; free input / free output per 1M tokens; provider documentation
  • Eden AI — model id flexai/deepseek-v4-flash-0731; $0.03 input / $0.1 output per 1M tokens; provider documentation
  • Vercel AI Gateway — model id deepseek/deepseek-v4-flash-0731; $0.076 input / $0.153 output per 1M tokens; provider documentation
  • OpenRouter — model id deepseek/deepseek-v4-flash-0731; $0.065 input / $0.18 output per 1M tokens; provider documentation
  • Ambient — model id deepseek/deepseek-v4-flash-0731; $0.08 input / $0.18 output per 1M tokens; provider documentation
  • Deep Infra — model id deepseek-ai/DeepSeek-V4-Flash-0731; $0.08 input / $0.18 output per 1M tokens; provider documentation
  • Eden AI — model id deepinfra/deepseek-ai/DeepSeek-V4-Flash-0731; $0.08 input / $0.18 output per 1M tokens; provider documentation
  • DigitalOcean — model id deepseek-v4-flash-0731; $0.08 input / $0.252 output per 1M tokens; provider documentation
  • Baseten — model id deepseek-ai/DeepSeek-V4-Flash-0731; $0.13 input / $0.26 output per 1M tokens; provider documentation
  • Perplexity Agent — model id deepseek/deepseek-v4-flash-0731; $0.13 input / $0.26 output per 1M tokens; provider documentation
  • RunInfra — model id deepseek-ai/DeepSeek-V4-Flash-0731; $0.13 input / $0.27 output per 1M tokens; provider documentation
  • Cortecs — model id deepseek-v4-flash-0731; $0.13 input / $0.28 output per 1M tokens; provider documentation
  • Inceptron — model id deepseek-ai/DeepSeek-V4-Flash-0731; $0.13 input / $0.28 output per 1M tokens; provider documentation
  • Weights & Biases — model id deepseek-ai/DeepSeek-V4-Flash-0731; $0.13 input / $0.28 output per 1M tokens; provider documentation
  • OpenReason — model id deepseek-ai/deepseek-v4-flash-0731; $0.137 input / $0.274 output per 1M tokens; provider documentation
  • AMD — model id DeepSeek-V4-Flash; $0.14 input / $0.28 output per 1M tokens; provider documentation
  • Eden AI — model id databricks/databricks-deepseek-v4-flash-0731; $0.14 input / $0.28 output per 1M tokens; provider documentation
  • Eden AI — model id nebius/deepseek-ai/DeepSeek-V4-Flash-0731; $0.14 input / $0.28 output per 1M tokens; provider documentation
  • Eden AI — model id together_ai/deepseek-ai/DeepSeek-V4-Flash-0731; $0.14 input / $0.28 output per 1M tokens; provider documentation
  • Hugging Face — model id deepseek-ai/DeepSeek-V4-Flash-0731; $0.14 input / $0.28 output per 1M tokens; provider documentation
  • NanoGPT — model id deepseek/deepseek-v4-flash-0731; $0.14 input / $0.28 output per 1M tokens; provider documentation
  • Nebius Token Factory — model id deepseek-ai/DeepSeek-V4-Flash-0731; $0.14 input / $0.28 output per 1M tokens; provider documentation
  • OpenRouter (batch) — model id deepseek/deepseek-v4-flash-0731:batch; $0.14 input / $0.28 output per 1M tokens; provider documentation
  • Requesty — model id deepseek-v4-flash-0731; $0.14 input / $0.28 output per 1M tokens; provider documentation
  • Requesty (EU) — model id deepseek-v4-flash-0731@eu; $0.14 input / $0.28 output per 1M tokens; provider documentation
  • Together AI — model id deepseek-ai/DeepSeek-V4-Flash-0731; $0.14 input / $0.28 output per 1M tokens; provider documentation
  • AIHubMix — model id deepseek-v4-flash-0731; $0.142 input / $0.284 output per 1M tokens; provider documentation
  • OrcaRouter — model id deepseek/deepseek-v4-flash-0731; $0.147 input / $0.295 output per 1M tokens; provider documentation
  • Venice AI — model id deepseek-v4-flash-0731; $0.175 input / $0.35 output per 1M tokens; provider documentation
  • Eden AI — model id tensorx/deepseek/deepseek-v4-flash-0731; $0.25 input / $0.3 output per 1M tokens; provider documentation
  • TensorX — model id deepseek/deepseek-v4-flash-0731; $0.25 input / $0.3 output per 1M tokens; provider documentation
  • GreenPT — model id deepseek-v4-flash-0731; $0.16 input / $0.399 output per 1M tokens; provider documentation
  • Alibaba — model id deepseek-v4-flash-0731; $0.2 input / $0.4 output per 1M tokens; provider documentation
  • AKI.IO — model id deepseek-v4-flash-0731-284b; $0.2 input / $0.5 output per 1M tokens; provider documentation
  • Eden AI — model id qwen/deepseek-v4-flash-0731; $0.176 input / $0.528 output per 1M tokens; provider documentation
  • Merge Gateway — model id deepseek/deepseek-v4-flash-0731-fast; $0.28 input / $0.56 output per 1M tokens; provider documentation
  • Eden AI — model id fireworks_ai/accounts/fireworks/models/deepseek-v4-flash-0731; $0.22 input / $0.66 output per 1M tokens; provider documentation
  • Fireworks AI — model id accounts/fireworks/models/deepseek-v4-flash-0731; $0.22 input / $0.66 output per 1M tokens; provider documentation
  • Merge Gateway — model id deepseek/deepseek-v4-flash-0731; $0.22 input / $0.66 output per 1M tokens; provider documentation
  • Tinfoil — model id deepseek-v4-flash; $0.3 input / $0.7 output per 1M tokens; provider documentation
  • Eden AI — model id scaleway/deepseek-v4-flash-0731; $0.465 input / $0.929 output per 1M tokens; provider documentation
  • Scaleway — model id deepseek-v4-flash-0731; $0.468 input / $0.936 output per 1M tokens; provider documentation
  • EmpirioLabs AI — model id deepseek-v4-flash-0731; $0.424 input / $1.27 output per 1M tokens; provider documentation
  • Charm Hyper — model id deepseek-v4-flash-0731; $0.44 input / $1.32 output per 1M tokens; provider documentation
  • Cloudflare Workers AI — model id @cf/deepseek-ai/deepseek-v4-flash-0731; $0.44 input / $1.32 output per 1M tokens; provider documentation
  • Eden AI — model id cloudflare/@cf/deepseek-ai/deepseek-v4-flash-0731; $0.44 input / $1.32 output per 1M tokens; provider documentation
  • Kilo Gateway — model id deepseek/deepseek-v4-flash-0731; $0.44 input / $1.32 output per 1M tokens; provider documentation
  • Ofox — model id deepseek/deepseek-v4-flash-0731; $0.44 input / $1.32 output per 1M tokens; provider documentation
  • Volcengine Ark — model id deepseek-v4-flash-ga-260731; $0.445 input / $1.34 output per 1M tokens; provider documentation
  • Ollama Cloud — model id deepseek-v4-flash:0731; not published input / not published output per 1M tokens; provider documentation