All AI models · LLM benchmarks · Methodology

Inception

Mercury 2.5 Preview

Mercury 2.5 is the fastest reasoning LLM, and the latest diffusion LLM (dLLM) from Inception. Instead of generating tokens sequentially, Mercury 2.5 produces and refines multiple tokens in parallel, achieving...

Model facts

Context window
26K tokens
Maximum output
66K tokens
Cheapest paid input
$0.04 per 1M tokens
Cheapest paid output
$0.15 per 1M tokens
Open weights
no
Providers
3

Capabilities and modalities

reasoning, tool calling, structured output, temperature control, text.

Observed price history

  • 2026-09-01: $0.04 input / $0.15 output per 1M tokens

API providers

  • NanoGPT — model id inception/mercury-2.5-preview; $0.04 input / $0.15 output per 1M tokens; provider documentation
  • OpenRouter — model id inception/mercury-2.5-preview; $0.04 input / $0.15 output per 1M tokens; provider documentation
  • Kilo Gateway (Inception) — model id inception/mercury-2.5-preview; $0.2 input / $0.75 output per 1M tokens; provider documentation