Thinking Machines
Inkling
Multimodal MoE reasoning model (975B total, 41B active) for text, image, and audio
Model facts
- Context window
- 1.05M tokens
- Maximum output
- 1.05M tokens
- Cheapest paid input
- $0.95 per 1M tokens
- Cheapest paid output
- $4.05 per 1M tokens
- Open weights
- yes
- Providers
- 25
Capabilities and modalities
reasoning, tool calling, structured output, attachments, temperature control, open weights, text, image, audio, pdf.
Local hardware estimate
At 4K context: Q4_K_M 583.9 GiB, Q8_0 1027.4 GiB, F16 1858.9 GiB working memory.
- Parameters
- 952.4 billion
- Architecture
- inkling
- Model license
- apache-2.0
Planning estimate adapted from the MIT-licensed whichllm estimator using metadata from Hugging Face; not a vendor minimum requirement.
Artificial Analysis benchmarks
Compare this model on the full LLM benchmark leaderboard.
| Configuration | AAI | Coding | Math | Output tokens/s |
|---|---|---|---|---|
| xhigh | 42.3 | 52.1 | not measured | 70.5 |
Observed price history
- 2026-08-24: $0.95 input / $4.05 output per 1M tokens
API providers
- Kilo Gateway (free) — model id thinkingmachines/inkling:free; free input / free output per 1M tokens; provider documentation
- Nvidia — model id thinkingmachines/inkling; free input / free output per 1M tokens; provider documentation
- OpenRouter (free) — model id thinkingmachines/inkling:free; free input / free output per 1M tokens; provider documentation
- Deep Infra — model id thinkingmachines/Inkling; $0.95 input / $4.05 output per 1M tokens; provider documentation
- Eden AI — model id deepinfra/thinkingmachines/Inkling; $0.95 input / $4.05 output per 1M tokens; provider documentation
- Kilo Gateway — model id thinkingmachines/inkling; $0.95 input / $4.05 output per 1M tokens; provider documentation
- Baseten — model id thinkingmachines/inkling; $1 input / $4.05 output per 1M tokens; provider documentation
- Eden AI — model id together_ai/thinkingmachines/Inkling; $1 input / $4.05 output per 1M tokens; provider documentation
- Fireworks AI — model id accounts/fireworks/models/inkling; $1 input / $4.05 output per 1M tokens; provider documentation
- Hugging Face — model id thinkingmachines/Inkling; $1 input / $4.05 output per 1M tokens; provider documentation
- Merge Gateway — model id thinkingmachines/inkling; $1 input / $4.05 output per 1M tokens; provider documentation
- NanoGPT — model id thinkingmachines/inkling; $1 input / $4.05 output per 1M tokens; provider documentation
- Neon — model id inkling; $1 input / $4.05 output per 1M tokens; provider documentation
- OpenRouter — model id thinkingmachines/inkling; $1 input / $4.05 output per 1M tokens; provider documentation
- OpenRouter (batch) — model id thinkingmachines/inkling:batch; $1 input / $4.05 output per 1M tokens; provider documentation
- Together AI — model id thinkingmachines/Inkling; $1 input / $4.05 output per 1M tokens; provider documentation
- Vercel AI Gateway — model id thinkingmachines/inkling; $1 input / $4.05 output per 1M tokens; provider documentation
- Modal — model id thinkingmachines/Inkling-NVFP4; $1.2 input / $5 output per 1M tokens; provider documentation
- Venice AI — model id inkling; $1.25 input / $5.06 output per 1M tokens; provider documentation
- Impossibl — model id thinkingmachines/inkling; $1.87 input / $4.68 output per 1M tokens; provider documentation
- LLMTR — model id thinkingmachines/inkling; $1.87 input / $4.68 output per 1M tokens; provider documentation
- Requesty — model id inkling; $1.87 input / $4.68 output per 1M tokens; provider documentation
- Thinking Machines — model id thinkingmachines/Inkling; $1.87 input / $4.68 output per 1M tokens; provider documentation
- Abacus — model id thinkingmachines/Inkling; $3.74 input / $9.36 output per 1M tokens; provider documentation
- Arena — model id inkling; not published input / not published output per 1M tokens; provider documentation