Inception
Mercury 2.5 Preview
Mercury 2.5 is the fastest reasoning LLM, and the latest diffusion LLM (dLLM) from Inception. Instead of generating tokens sequentially, Mercury 2.5 produces and refines multiple tokens in parallel, achieving...
Model facts
- Context window
- 26K tokens
- Maximum output
- 66K tokens
- Cheapest paid input
- $0.04 per 1M tokens
- Cheapest paid output
- $0.15 per 1M tokens
- Open weights
- no
- Providers
- 3
Capabilities and modalities
reasoning, tool calling, structured output, temperature control, text.
Observed price history
- 2026-09-01: $0.04 input / $0.15 output per 1M tokens
API providers
- NanoGPT — model id inception/mercury-2.5-preview; $0.04 input / $0.15 output per 1M tokens; provider documentation
- OpenRouter — model id inception/mercury-2.5-preview; $0.04 input / $0.15 output per 1M tokens; provider documentation
- Kilo Gateway (Inception) — model id inception/mercury-2.5-preview; $0.2 input / $0.75 output per 1M tokens; provider documentation