NVIDIA
llama-3.1-nemotron-ultra-253b-v1
A reasoning-optimized LLM based on Llama 3.1, Nemotron Ultra 253B delivers strong performance in tasks like RAG and tool use, with high efficiency and reduced latency.
Model facts
- Context window
- 128K tokens
- Maximum output
- 128K tokens
- Cheapest paid input
- $0.598 per 1M tokens
- Cheapest paid output
- $1.79 per 1M tokens
- Open weights
- no
- Providers
- 1
Capabilities and modalities
reasoning, tool calling, structured output, text.
Artificial Analysis benchmarks
Compare this model on the full LLM benchmark leaderboard.
| Configuration | AAI | Coding | Math | Output tokens/s |
|---|---|---|---|---|
| Reasoning | 8.9 | not measured | 63.7 | 52.9 |
Observed price history
- 2026-08-24: $0.598 input / $1.79 output per 1M tokens
API providers
- Cortecs — model id llama-3.1-nemotron-ultra-253b-v1; $0.598 input / $1.79 output per 1M tokens; provider documentation