NVIDIA
Nemotron 3.5 Lightning 30B A3B
Fast NVIDIA Nemotron MoE for reliable agentic tasks across enterprise workloads
Model facts
- Context window
- 262K tokens
- Maximum output
- 262K tokens
- Cheapest paid input
- $0.05 per 1M tokens
- Cheapest paid output
- $0.15 per 1M tokens
- Open weights
- yes
- Providers
- 10
Capabilities and modalities
reasoning, tool calling, structured output, temperature control, open weights, text.
Local hardware estimate
At 4K context: Q4_K_M 17.9 GiB, Q8_0 32.6 GiB, F16 60.2 GiB working memory.
- Parameters
- 31.6 billion
- Architecture
- nemotronh
- Model license
- other
Planning estimate adapted from the MIT-licensed whichllm estimator using metadata from Hugging Face; not a vendor minimum requirement.
Observed price history
- 2026-08-24: $0.05 input / $0.15 output per 1M tokens
API providers
- Merge Gateway — model id nvidia/nemotron-3.5-lightning-30b-a3b; free input / free output per 1M tokens; provider documentation
- Nvidia — model id nvidia/nemotron-3.5-lightning-30b-a3b; free input / free output per 1M tokens; provider documentation
- Requesty — model id nemotron-3.5-lightning-30b-a3b; free input / free output per 1M tokens; provider documentation
- RunInfra — model id nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16; $0.05 input / $0.15 output per 1M tokens; provider documentation
- Fireworks AI — model id accounts/fireworks/models/nemotron-lightning-3p5-30b-a3b; $0.05 input / $0.2 output per 1M tokens; provider documentation
- Requesty — model id nemotron-lightning-3.5-30b-a3b; $0.05 input / $0.2 output per 1M tokens; provider documentation
- OpenRouter — model id nvidia/nemotron-3.5-lightning; $0.08 input / $0.2 output per 1M tokens; provider documentation
- Kilo Gateway — model id nvidia/nemotron-3.5-lightning; $0.08 input / $0.2 output per 1M tokens; provider documentation
- Nebius Token Factory — model id nvidia/Nemotron-3_5-Lightning; $0.06 input / $0.24 output per 1M tokens; provider documentation
- Pioneer — model id nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16; $0.5 input / $0.5 output per 1M tokens; provider documentation