DeepSeek
DeepSeek V4 Flash
Fast DeepSeek V4 lane for economical reasoning, coding, and long-context work
Model facts
- Context window
- 1M tokens
- Maximum output
- 384K tokens
- Cheapest paid input
- $0.051 per 1M tokens
- Cheapest paid output
- $0.104 per 1M tokens
- Open weights
- yes
- Providers
- 81
Capabilities and modalities
reasoning, tool calling, structured output, temperature control, open weights, text.
Local hardware estimate
At 4K context: Q4_K_M 179.0 GiB, Q8_0 314.5 GiB, F16 568.6 GiB working memory.
- Parameters
- 290.9 billion
- Architecture
- deepseekv4
- Model license
- mit
Planning estimate adapted from the MIT-licensed whichllm estimator using metadata from Hugging Face; not a vendor minimum requirement.
Artificial Analysis benchmarks
Compare this model on the full LLM benchmark leaderboard.
| Configuration | AAI | Coding | Math | Output tokens/s |
|---|---|---|---|---|
| Reasoning, Max Effort | 42.1 | 56.2 | not measured | not measured |
| Reasoning, High Effort | 39.0 | 52.0 | not measured | not measured |
| Non-reasoning | 29.3 | not measured | not measured | not measured |
Observed price history
- 2026-08-24: $0.056 input / $0.112 output per 1M tokens
- 2026-08-25: $0.062 input / $0.125 output per 1M tokens
- 2026-08-26: $0.051 input / $0.104 output per 1M tokens
API providers
- Alibaba Token Plan — model id deepseek-v4-flash; free input / free output per 1M tokens; provider documentation
- Alibaba Token Plan (China) — model id deepseek-v4-flash; free input / free output per 1M tokens; provider documentation
- InferX — model id deepseek-v4-flash; free input / free output per 1M tokens; provider documentation
- Kenari — model id deepseek-v4-flash; free input / free output per 1M tokens; provider documentation
- Kenari (Free) — model id deepseek-v4-flash:free; free input / free output per 1M tokens; provider documentation
- NaN — model id deepseek-v4-flash; free input / free output per 1M tokens; provider documentation
- OrcaRouter (free) — model id deepseek/deepseek-v4-flash-free; free input / free output per 1M tokens; provider documentation
- Pendra — model id deepseek-v4-flash; free input / free output per 1M tokens; provider documentation
- SCNet Token Plan — model id DeepSeek-V4-Flash; free input / free output per 1M tokens; provider documentation
- SenseNova (China) — model id deepseek-v4-flash; free input / free output per 1M tokens; provider documentation
- Umans AI Coding Plan — model id umans-deepseek-v4-flash-0731; free input / free output per 1M tokens; provider documentation
- UnoRouter — model id deepseek-v4-flash:free; free input / free output per 1M tokens; provider documentation
- Volcengine Ark Coding Plan — model id deepseek-v4-flash; free input / free output per 1M tokens; provider documentation
- DevPass (LLM Gateway) — model id deepseek-v4-flash; $0.051 input / $0.104 output per 1M tokens; provider documentation
- LLM Gateway (Gonka24) — model id gonka24/deepseek-v4-flash; $0.051 input / $0.104 output per 1M tokens; provider documentation
- CrofAI (New) — model id deepseek-v4-flash-0731; $0.08 input / $0.1 output per 1M tokens; provider documentation
- UnoRouter — model id deepseek-v4-flash; $0.062 input / $0.125 output per 1M tokens; provider documentation
- LLM Gateway (Runware) — model id runware/deepseek-v4-flash; $0.076 input / $0.153 output per 1M tokens; provider documentation
- DigitalOcean — model id deepseek-4-flash; $0.068 input / $0.168 output per 1M tokens; provider documentation
- LLM Gateway (DeepInfra) — model id deepinfra/deepseek-v4-flash; $0.08 input / $0.18 output per 1M tokens; provider documentation
- OpenRouter — model id deepseek/deepseek-v4-flash; $0.089 input / $0.177 output per 1M tokens; provider documentation
- Deep Infra — model id deepseek-ai/DeepSeek-V4-Flash; $0.09 input / $0.18 output per 1M tokens; provider documentation
- TokenGo — model id deepseek/deepseek-v4-flash; $0.098 input / $0.196 output per 1M tokens; provider documentation
- Modelis — model id deepseek-v4-flash; $0.098 input / $0.197 output per 1M tokens; provider documentation
- Pioneer — model id deepseek-ai/DeepSeek-V4-Flash; $0.1 input / $0.2 output per 1M tokens; provider documentation
- CrofAI — model id deepseek-v4-flash; $0.12 input / $0.21 output per 1M tokens; provider documentation
- GMI Cloud — model id deepseek-ai/DeepSeek-V4-Flash; $0.112 input / $0.224 output per 1M tokens; provider documentation
- routing.run — model id deepseek-v4-flash; $0.112 input / $0.224 output per 1M tokens; provider documentation
- Vercel AI Gateway — model id deepseek/deepseek-v4-flash; $0.13 input / $0.26 output per 1M tokens; provider documentation
- LLM Gateway (Consensus Protocol) — model id consensusprotocol/deepseek-v4-flash; $0.13 input / $0.27 output per 1M tokens; provider documentation
- ai& — model id deepseek-ai/deepseek-v4-flash; $0.15 input / $0.25 output per 1M tokens; provider documentation
- AIHubMix (Alibaba Cloud) — model id alicloud-deepseek-v4-flash; $0.14 input / $0.28 output per 1M tokens; provider documentation
- Abacus — model id deepseek-ai/DeepSeek-V4-Flash; $0.14 input / $0.28 output per 1M tokens; provider documentation
- Alibaba (China) — model id deepseek-v4-flash; $0.14 input / $0.28 output per 1M tokens; provider documentation
- Ambient — model id deepseek/deepseek-v4-flash; $0.14 input / $0.28 output per 1M tokens; provider documentation
- Auriko — model id deepseek-v4-flash; $0.14 input / $0.28 output per 1M tokens; provider documentation
- ClinePass — model id cline-pass/deepseek-v4-flash; $0.14 input / $0.28 output per 1M tokens; provider documentation
- DeepSeek — model id deepseek-v4-flash; $0.14 input / $0.28 output per 1M tokens; provider documentation
- EmpirioLabs AI — model id deepseek-v4-flash; $0.14 input / $0.28 output per 1M tokens; provider documentation
- HPC-AI — model id deepseek/deepseek-v4-flash; $0.14 input / $0.28 output per 1M tokens; provider documentation
- Hugging Face — model id deepseek-ai/DeepSeek-V4-Flash; $0.14 input / $0.28 output per 1M tokens; provider documentation
- Jalapeno Cloud — model id DeepSeek-V4-Flash; $0.14 input / $0.28 output per 1M tokens; provider documentation
- Kilo Gateway — model id deepseek/deepseek-v4-flash; $0.14 input / $0.28 output per 1M tokens; provider documentation
- LLM Gateway (CanopyWave) — model id canopywave/deepseek-v4-flash; $0.14 input / $0.28 output per 1M tokens; provider documentation
- LLM Gateway (DeepSeek) — model id deepseek/deepseek-v4-flash; $0.14 input / $0.28 output per 1M tokens; provider documentation
- LLM Gateway (NovitaAI) — model id novita/deepseek-v4-flash; $0.14 input / $0.28 output per 1M tokens; provider documentation
- LLM Gateway (RanoAI) — model id ranoai/deepseek-v4-flash; $0.14 input / $0.28 output per 1M tokens; provider documentation
- LLM Gateway (Together AI) — model id together-ai/deepseek-v4-flash; $0.14 input / $0.28 output per 1M tokens; provider documentation
- NanoGPT — model id deepseek/deepseek-v4-flash; $0.14 input / $0.28 output per 1M tokens; provider documentation
- Neuralwatt — model id deepseek-v4-flash; $0.14 input / $0.28 output per 1M tokens; provider documentation
- NovitaAI — model id deepseek/deepseek-v4-flash; $0.14 input / $0.28 output per 1M tokens; provider documentation
- Nvidia — model id deepseek-ai/deepseek-v4-flash; $0.14 input / $0.28 output per 1M tokens; provider documentation
- OpenCode Zen — model id deepseek-v4-flash; $0.14 input / $0.28 output per 1M tokens; provider documentation
- Requesty — model id deepseek-v4-flash; $0.14 input / $0.28 output per 1M tokens; provider documentation
- SiliconFlow — model id deepseek-ai/DeepSeek-V4-Flash; $0.14 input / $0.28 output per 1M tokens; provider documentation
- SiliconFlow (China) — model id deepseek-ai/DeepSeek-V4-Flash; $0.14 input / $0.28 output per 1M tokens; provider documentation
- Umans AI — model id umans-deepseek-v4-flash-0731; $0.14 input / $0.28 output per 1M tokens; provider documentation
- Weights & Biases — model id deepseek-ai/DeepSeek-V4-Flash; $0.14 input / $0.28 output per 1M tokens; provider documentation
- ZenMux — model id deepseek/deepseek-v4-flash; $0.14 input / $0.28 output per 1M tokens; provider documentation
- EBCloud — model id DeepSeek-V4-Flash; $0.143 input / $0.286 output per 1M tokens; provider documentation
- OrcaRouter — model id deepseek/deepseek-v4-flash; $0.147 input / $0.295 output per 1M tokens; provider documentation
- TensorX — model id deepseek/deepseek-v4-flash; $0.15 input / $0.3 output per 1M tokens; provider documentation
- Vivgrid — model id deepseek-v4-flash; $0.15 input / $0.3 output per 1M tokens; provider documentation
- AIHubMix (DeepSeek) — model id deep-deepseek-v4-flash; $0.154 input / $0.308 output per 1M tokens; provider documentation
- Charm Hyper — model id deepseek-v4-flash; $0.2 input / $0.4 output per 1M tokens; provider documentation
- LLM Gateway (Alibaba Cloud) — model id alibaba/deepseek-v4-flash; $0.2 input / $0.4 output per 1M tokens; provider documentation
- Azure — model id deepseek-v4-flash; $0.19 input / $0.51 output per 1M tokens; provider documentation
- Impossibl — model id deepseek/deepseek-v4-flash; $0.19 input / $0.51 output per 1M tokens; provider documentation
- LLM Gateway (Fireworks AI) — model id fireworks/deepseek-v4-flash; $0.22 input / $0.66 output per 1M tokens; provider documentation
- Merge Gateway — model id deepseek/deepseek-v4-flash; $0.22 input / $0.66 output per 1M tokens; provider documentation
- OpenCode Go — model id deepseek-v4-flash; $0.22 input / $0.66 output per 1M tokens; provider documentation
- Vancine — model id deepseek-v4-flash; $0.22 input / $0.66 output per 1M tokens; provider documentation
- above.dev — model id deepseek-v4-flash; $0.242 input / $0.726 output per 1M tokens; provider documentation
- Vultr — model id deepseek-ai/DeepSeek-V4-Flash; $0.3 input / $1 output per 1M tokens; provider documentation
- CrossModel — model id deepseek/deepseek-v4-flash; $0.405 input / $1.22 output per 1M tokens; provider documentation
- Eden AI — model id deepseek/deepseek-v4-flash; $0.44 input / $1.32 output per 1M tokens; provider documentation
- LLM Gateway (Baidu) — model id baidu/deepseek-v4-flash; $0.44 input / $1.32 output per 1M tokens; provider documentation
- LLM Gateway (ByteDance) — model id bytedance/deepseek-v4-flash; $0.44 input / $1.32 output per 1M tokens; provider documentation
- Ofox — model id deepseek/deepseek-v4-flash; $0.44 input / $1.32 output per 1M tokens; provider documentation
- AnyAPI — model id deepseek/deepseek-v4-flash; not published input / not published output per 1M tokens; provider documentation
- Ollama Cloud — model id deepseek-v4-flash; not published input / not published output per 1M tokens; provider documentation