NVIDIA
Nemotron Nano 9B V2
Compact Nemotron model for efficient reasoning and deployable AI agents
Model facts
- Context window
- 131K tokens
- Maximum output
- 131K tokens
- Cheapest paid input
- $0.06 per 1M tokens
- Cheapest paid output
- $0.23 per 1M tokens
- Open weights
- yes
- Providers
- 3
Capabilities and modalities
reasoning, tool calling, structured output, temperature control, open weights, text.
Local hardware estimate
At 4K context: Q4_K_M 6.4 GiB, Q8_0 10.6 GiB, F16 18.3 GiB working memory.
- Parameters
- 8.9 billion
- Architecture
- nemotronh
- Model license
- other
Planning estimate adapted from the MIT-licensed whichllm estimator using metadata from Hugging Face; not a vendor minimum requirement.
Artificial Analysis benchmarks
Compare this model on the full LLM benchmark leaderboard.
| Configuration | AAI | Coding | Math | Output tokens/s |
|---|---|---|---|---|
| Reasoning | 8.7 | not measured | 69.7 | 100.5 |
| Non-reasoning | 7.2 | not measured | 62.3 | 163.6 |
Observed price history
- 2026-08-24: $0.06 input / $0.23 output per 1M tokens
API providers
- Nvidia — model id nvidia/nvidia-nemotron-nano-9b-v2; free input / free output per 1M tokens; provider documentation
- Amazon Bedrock (NVIDIA) — model id nvidia.nemotron-nano-9b-v2; $0.06 input / $0.23 output per 1M tokens; provider documentation
- Vercel AI Gateway (Nvidia) — model id nvidia/nemotron-nano-9b-v2; $0.06 input / $0.23 output per 1M tokens; provider documentation