All AI models · LLM benchmarks · Methodology

NVIDIA

Nemotron Nano 9B V2

Compact Nemotron model for efficient reasoning and deployable AI agents

Model facts

Context window
131K tokens
Maximum output
131K tokens
Cheapest paid input
$0.06 per 1M tokens
Cheapest paid output
$0.23 per 1M tokens
Open weights
yes
Providers
3

Capabilities and modalities

reasoning, tool calling, structured output, temperature control, open weights, text.

Local hardware estimate

At 4K context: Q4_K_M 6.4 GiB, Q8_0 10.6 GiB, F16 18.3 GiB working memory.

Parameters
8.9 billion
Architecture
nemotronh
Model license
other

Planning estimate adapted from the MIT-licensed whichllm estimator using metadata from Hugging Face; not a vendor minimum requirement.

Artificial Analysis benchmarks

Compare this model on the full LLM benchmark leaderboard.

ConfigurationAAICodingMathOutput tokens/s
Reasoning8.7not measured69.7100.5
Non-reasoning7.2not measured62.3163.6

Observed price history

  • 2026-08-24: $0.06 input / $0.23 output per 1M tokens

API providers

  • Nvidia — model id nvidia/nvidia-nemotron-nano-9b-v2; free input / free output per 1M tokens; provider documentation
  • Amazon Bedrock (NVIDIA) — model id nvidia.nemotron-nano-9b-v2; $0.06 input / $0.23 output per 1M tokens; provider documentation
  • Vercel AI Gateway (Nvidia) — model id nvidia/nemotron-nano-9b-v2; $0.06 input / $0.23 output per 1M tokens; provider documentation