All AI models · LLM benchmarks · Methodology

NVIDIA

llama-3.1-nemotron-ultra-253b-v1

A reasoning-optimized LLM based on Llama 3.1, Nemotron Ultra 253B delivers strong performance in tasks like RAG and tool use, with high efficiency and reduced latency.

Model facts

Context window
128K tokens
Maximum output
128K tokens
Cheapest paid input
$0.598 per 1M tokens
Cheapest paid output
$1.79 per 1M tokens
Open weights
no
Providers
1

Capabilities and modalities

reasoning, tool calling, structured output, text.

Artificial Analysis benchmarks

Compare this model on the full LLM benchmark leaderboard.

ConfigurationAAICodingMathOutput tokens/s
Reasoning8.9not measured63.752.9

Observed price history

  • 2026-08-24: $0.598 input / $1.79 output per 1M tokens

API providers

  • Cortecs — model id llama-3.1-nemotron-ultra-253b-v1; $0.598 input / $1.79 output per 1M tokens; provider documentation