All AI models · LLM benchmarks · Methodology

NVIDIA

NVIDIA: Nemotron 3 Ultra

NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built on a hybrid Transformer-Mamba mixture-of-experts architecture, it... **Terms of service** For NVIDIA free endpoints (Super/Ultra/etc): Trial use only - do not submit personal or confidential data. Your use is logged for security purposes and to improve NVIDIA products and services. The logged session data for improvement purposes is not linked to your identity or any persistent identifier. For more information about our data processing practices, see our [Privacy Policy](https://www.nvidia.com/en-us/about-nvidia/privacy-policy/). By interacting with this endpoint, you consent to our collection, recording, and use of such information and the [NVIDIA API Trial Terms of Service](https://assets.ngc.nvidia.com/products/api-catalog/legal/NVIDIA%20API%20Trial%20Terms%20of%20Service.pdf).

Model facts

Context window
1M tokens
Maximum output
66K tokens
Cheapest paid input
free per 1M tokens
Cheapest paid output
free per 1M tokens
Open weights
yes
Providers
1

Capabilities and modalities

reasoning, tool calling, temperature control, open weights, text.

Observed price history

  • 2026-08-24: free input / free output per 1M tokens

API providers

  • Kilo Gateway (free) — model id nvidia/nemotron-3-ultra-550b-a55b:free; free input / free output per 1M tokens; provider documentation