NVIDIA
Nemotron 3 Ultra
Largest Nemotron 3 model for maximum open-weight reasoning and agent accuracy
Model facts
- Context window
- 1M tokens
- Maximum output
- 131K tokens
- Cheapest paid input
- $0.9 per 1M tokens
- Cheapest paid output
- $1.7 per 1M tokens
- Open weights
- yes
- Providers
- 8
Capabilities and modalities
reasoning, tool calling, structured output, temperature control, open weights, text.
Observed price history
- 2026-08-24: $0.9 input / $1.7 output per 1M tokens
API providers
- Bothub (free) — model id nemotron-3-ultra-550b-a55b:free; free input / free output per 1M tokens; provider documentation
- OpenRouter (free) — model id nvidia/nemotron-3-ultra-550b-a55b:free; free input / free output per 1M tokens; provider documentation
- DigitalOcean — model id nemotron-3-ultra-550b; $0.9 input / $1.7 output per 1M tokens; provider documentation
- Requesty — model id nvidia-nemotron-3-ultra; $0.5 input / $2.5 output per 1M tokens; provider documentation
- Vercel AI Gateway — model id nvidia/nemotron-3-ultra-550b-a55b; $0.6 input / $2.4 output per 1M tokens; provider documentation
- Weights & Biases — model id nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B; $0.75 input / $2.75 output per 1M tokens; provider documentation
- Venice AI (NVIDIA) — model id nvidia-nemotron-3-ultra-550b-a55b; $0.625 input / $3.12 output per 1M tokens; provider documentation
- Ollama Cloud — model id nemotron-3-ultra; not published input / not published output per 1M tokens; provider documentation