All AI models · LLM benchmarks · Methodology

DeepSeek

DeepSeek V4 Flash

Fast DeepSeek V4 lane for economical reasoning, coding, and long-context work

Model facts

Context window
1M tokens
Maximum output
384K tokens
Cheapest paid input
$0.051 per 1M tokens
Cheapest paid output
$0.104 per 1M tokens
Open weights
yes
Providers
81

Capabilities and modalities

reasoning, tool calling, structured output, temperature control, open weights, text.

Local hardware estimate

At 4K context: Q4_K_M 179.0 GiB, Q8_0 314.5 GiB, F16 568.6 GiB working memory.

Parameters
290.9 billion
Architecture
deepseekv4
Model license
mit

Planning estimate adapted from the MIT-licensed whichllm estimator using metadata from Hugging Face; not a vendor minimum requirement.

Artificial Analysis benchmarks

Compare this model on the full LLM benchmark leaderboard.

ConfigurationAAICodingMathOutput tokens/s
Reasoning, Max Effort42.156.2not measurednot measured
Reasoning, High Effort39.052.0not measurednot measured
Non-reasoning29.3not measurednot measurednot measured

Observed price history

  • 2026-08-24: $0.056 input / $0.112 output per 1M tokens
  • 2026-08-25: $0.062 input / $0.125 output per 1M tokens
  • 2026-08-26: $0.051 input / $0.104 output per 1M tokens

API providers

  • Alibaba Token Plan — model id deepseek-v4-flash; free input / free output per 1M tokens; provider documentation
  • Alibaba Token Plan (China) — model id deepseek-v4-flash; free input / free output per 1M tokens; provider documentation
  • InferX — model id deepseek-v4-flash; free input / free output per 1M tokens; provider documentation
  • Kenari — model id deepseek-v4-flash; free input / free output per 1M tokens; provider documentation
  • Kenari (Free) — model id deepseek-v4-flash:free; free input / free output per 1M tokens; provider documentation
  • NaN — model id deepseek-v4-flash; free input / free output per 1M tokens; provider documentation
  • OrcaRouter (free) — model id deepseek/deepseek-v4-flash-free; free input / free output per 1M tokens; provider documentation
  • Pendra — model id deepseek-v4-flash; free input / free output per 1M tokens; provider documentation
  • SCNet Token Plan — model id DeepSeek-V4-Flash; free input / free output per 1M tokens; provider documentation
  • SenseNova (China) — model id deepseek-v4-flash; free input / free output per 1M tokens; provider documentation
  • Umans AI Coding Plan — model id umans-deepseek-v4-flash-0731; free input / free output per 1M tokens; provider documentation
  • UnoRouter — model id deepseek-v4-flash:free; free input / free output per 1M tokens; provider documentation
  • Volcengine Ark Coding Plan — model id deepseek-v4-flash; free input / free output per 1M tokens; provider documentation
  • DevPass (LLM Gateway) — model id deepseek-v4-flash; $0.051 input / $0.104 output per 1M tokens; provider documentation
  • LLM Gateway (Gonka24) — model id gonka24/deepseek-v4-flash; $0.051 input / $0.104 output per 1M tokens; provider documentation
  • CrofAI (New) — model id deepseek-v4-flash-0731; $0.08 input / $0.1 output per 1M tokens; provider documentation
  • UnoRouter — model id deepseek-v4-flash; $0.062 input / $0.125 output per 1M tokens; provider documentation
  • LLM Gateway (Runware) — model id runware/deepseek-v4-flash; $0.076 input / $0.153 output per 1M tokens; provider documentation
  • DigitalOcean — model id deepseek-4-flash; $0.068 input / $0.168 output per 1M tokens; provider documentation
  • LLM Gateway (DeepInfra) — model id deepinfra/deepseek-v4-flash; $0.08 input / $0.18 output per 1M tokens; provider documentation
  • OpenRouter — model id deepseek/deepseek-v4-flash; $0.089 input / $0.177 output per 1M tokens; provider documentation
  • Deep Infra — model id deepseek-ai/DeepSeek-V4-Flash; $0.09 input / $0.18 output per 1M tokens; provider documentation
  • TokenGo — model id deepseek/deepseek-v4-flash; $0.098 input / $0.196 output per 1M tokens; provider documentation
  • Modelis — model id deepseek-v4-flash; $0.098 input / $0.197 output per 1M tokens; provider documentation
  • Pioneer — model id deepseek-ai/DeepSeek-V4-Flash; $0.1 input / $0.2 output per 1M tokens; provider documentation
  • CrofAI — model id deepseek-v4-flash; $0.12 input / $0.21 output per 1M tokens; provider documentation
  • GMI Cloud — model id deepseek-ai/DeepSeek-V4-Flash; $0.112 input / $0.224 output per 1M tokens; provider documentation
  • routing.run — model id deepseek-v4-flash; $0.112 input / $0.224 output per 1M tokens; provider documentation
  • Vercel AI Gateway — model id deepseek/deepseek-v4-flash; $0.13 input / $0.26 output per 1M tokens; provider documentation
  • LLM Gateway (Consensus Protocol) — model id consensusprotocol/deepseek-v4-flash; $0.13 input / $0.27 output per 1M tokens; provider documentation
  • ai& — model id deepseek-ai/deepseek-v4-flash; $0.15 input / $0.25 output per 1M tokens; provider documentation
  • AIHubMix (Alibaba Cloud) — model id alicloud-deepseek-v4-flash; $0.14 input / $0.28 output per 1M tokens; provider documentation
  • Abacus — model id deepseek-ai/DeepSeek-V4-Flash; $0.14 input / $0.28 output per 1M tokens; provider documentation
  • Alibaba (China) — model id deepseek-v4-flash; $0.14 input / $0.28 output per 1M tokens; provider documentation
  • Ambient — model id deepseek/deepseek-v4-flash; $0.14 input / $0.28 output per 1M tokens; provider documentation
  • Auriko — model id deepseek-v4-flash; $0.14 input / $0.28 output per 1M tokens; provider documentation
  • ClinePass — model id cline-pass/deepseek-v4-flash; $0.14 input / $0.28 output per 1M tokens; provider documentation
  • DeepSeek — model id deepseek-v4-flash; $0.14 input / $0.28 output per 1M tokens; provider documentation
  • EmpirioLabs AI — model id deepseek-v4-flash; $0.14 input / $0.28 output per 1M tokens; provider documentation
  • HPC-AI — model id deepseek/deepseek-v4-flash; $0.14 input / $0.28 output per 1M tokens; provider documentation
  • Hugging Face — model id deepseek-ai/DeepSeek-V4-Flash; $0.14 input / $0.28 output per 1M tokens; provider documentation
  • Jalapeno Cloud — model id DeepSeek-V4-Flash; $0.14 input / $0.28 output per 1M tokens; provider documentation
  • Kilo Gateway — model id deepseek/deepseek-v4-flash; $0.14 input / $0.28 output per 1M tokens; provider documentation
  • LLM Gateway (CanopyWave) — model id canopywave/deepseek-v4-flash; $0.14 input / $0.28 output per 1M tokens; provider documentation
  • LLM Gateway (DeepSeek) — model id deepseek/deepseek-v4-flash; $0.14 input / $0.28 output per 1M tokens; provider documentation
  • LLM Gateway (NovitaAI) — model id novita/deepseek-v4-flash; $0.14 input / $0.28 output per 1M tokens; provider documentation
  • LLM Gateway (RanoAI) — model id ranoai/deepseek-v4-flash; $0.14 input / $0.28 output per 1M tokens; provider documentation
  • LLM Gateway (Together AI) — model id together-ai/deepseek-v4-flash; $0.14 input / $0.28 output per 1M tokens; provider documentation
  • NanoGPT — model id deepseek/deepseek-v4-flash; $0.14 input / $0.28 output per 1M tokens; provider documentation
  • Neuralwatt — model id deepseek-v4-flash; $0.14 input / $0.28 output per 1M tokens; provider documentation
  • NovitaAI — model id deepseek/deepseek-v4-flash; $0.14 input / $0.28 output per 1M tokens; provider documentation
  • Nvidia — model id deepseek-ai/deepseek-v4-flash; $0.14 input / $0.28 output per 1M tokens; provider documentation
  • OpenCode Zen — model id deepseek-v4-flash; $0.14 input / $0.28 output per 1M tokens; provider documentation
  • Requesty — model id deepseek-v4-flash; $0.14 input / $0.28 output per 1M tokens; provider documentation
  • SiliconFlow — model id deepseek-ai/DeepSeek-V4-Flash; $0.14 input / $0.28 output per 1M tokens; provider documentation
  • SiliconFlow (China) — model id deepseek-ai/DeepSeek-V4-Flash; $0.14 input / $0.28 output per 1M tokens; provider documentation
  • Umans AI — model id umans-deepseek-v4-flash-0731; $0.14 input / $0.28 output per 1M tokens; provider documentation
  • Weights & Biases — model id deepseek-ai/DeepSeek-V4-Flash; $0.14 input / $0.28 output per 1M tokens; provider documentation
  • ZenMux — model id deepseek/deepseek-v4-flash; $0.14 input / $0.28 output per 1M tokens; provider documentation
  • EBCloud — model id DeepSeek-V4-Flash; $0.143 input / $0.286 output per 1M tokens; provider documentation
  • OrcaRouter — model id deepseek/deepseek-v4-flash; $0.147 input / $0.295 output per 1M tokens; provider documentation
  • TensorX — model id deepseek/deepseek-v4-flash; $0.15 input / $0.3 output per 1M tokens; provider documentation
  • Vivgrid — model id deepseek-v4-flash; $0.15 input / $0.3 output per 1M tokens; provider documentation
  • AIHubMix (DeepSeek) — model id deep-deepseek-v4-flash; $0.154 input / $0.308 output per 1M tokens; provider documentation
  • Charm Hyper — model id deepseek-v4-flash; $0.2 input / $0.4 output per 1M tokens; provider documentation
  • LLM Gateway (Alibaba Cloud) — model id alibaba/deepseek-v4-flash; $0.2 input / $0.4 output per 1M tokens; provider documentation
  • Azure — model id deepseek-v4-flash; $0.19 input / $0.51 output per 1M tokens; provider documentation
  • Impossibl — model id deepseek/deepseek-v4-flash; $0.19 input / $0.51 output per 1M tokens; provider documentation
  • LLM Gateway (Fireworks AI) — model id fireworks/deepseek-v4-flash; $0.22 input / $0.66 output per 1M tokens; provider documentation
  • Merge Gateway — model id deepseek/deepseek-v4-flash; $0.22 input / $0.66 output per 1M tokens; provider documentation
  • OpenCode Go — model id deepseek-v4-flash; $0.22 input / $0.66 output per 1M tokens; provider documentation
  • Vancine — model id deepseek-v4-flash; $0.22 input / $0.66 output per 1M tokens; provider documentation
  • above.dev — model id deepseek-v4-flash; $0.242 input / $0.726 output per 1M tokens; provider documentation
  • Vultr — model id deepseek-ai/DeepSeek-V4-Flash; $0.3 input / $1 output per 1M tokens; provider documentation
  • CrossModel — model id deepseek/deepseek-v4-flash; $0.405 input / $1.22 output per 1M tokens; provider documentation
  • Eden AI — model id deepseek/deepseek-v4-flash; $0.44 input / $1.32 output per 1M tokens; provider documentation
  • LLM Gateway (Baidu) — model id baidu/deepseek-v4-flash; $0.44 input / $1.32 output per 1M tokens; provider documentation
  • LLM Gateway (ByteDance) — model id bytedance/deepseek-v4-flash; $0.44 input / $1.32 output per 1M tokens; provider documentation
  • Ofox — model id deepseek/deepseek-v4-flash; $0.44 input / $1.32 output per 1M tokens; provider documentation
  • AnyAPI — model id deepseek/deepseek-v4-flash; not published input / not published output per 1M tokens; provider documentation
  • Ollama Cloud — model id deepseek-v4-flash; not published input / not published output per 1M tokens; provider documentation