InclusionAI
Ling 3.0 Flash
*Ling-3.0-flash* is a *124B-parameter Mixture-of-Experts (MoE) model*, with approximately *5.1B parameters activated per token*. The model is designed with *token efficiency and production-scale agentic inference* as key priorities, enabling developers...
Model facts
- Context window
- 262K tokens
- Maximum output
- 33K tokens
- Cheapest paid input
- $0.021 per 1M tokens
- Cheapest paid output
- $0.063 per 1M tokens
- Open weights
- yes
- Providers
- 5
Capabilities and modalities
reasoning, tool calling, temperature control, open weights, text.
Local hardware estimate
At 4K context: Q4_K_M 74.5 GiB, Q8_0 133.9 GiB, F16 245.2 GiB working memory.
- Parameters
- 127.5 billion
- Architecture
- bailingmoev3
- Model license
- mit
Planning estimate adapted from the MIT-licensed whichllm estimator using metadata from Hugging Face; not a vendor minimum requirement.
Artificial Analysis benchmarks
Compare this model on the full LLM benchmark leaderboard.
| Configuration | AAI | Coding | Math | Output tokens/s |
|---|---|---|---|---|
| default | 37.8 | 50.6 | not measured | 336.0 |
Observed price history
- 2026-08-24: $0.021 input / $0.063 output per 1M tokens
API providers
- OpenRouter — model id inclusionai/ling-3.0-flash; $0.021 input / $0.063 output per 1M tokens; provider documentation
- DevPass (LLM Gateway) (InclusionAI) — model id ling-3.0-flash; $0.06 input / $0.18 output per 1M tokens; provider documentation
- Kilo Gateway — model id inclusionai/ling-3.0-flash; $0.06 input / $0.18 output per 1M tokens; provider documentation
- Vercel AI Gateway — model id inclusionai/ling-3.0-flash; $0.06 input / $0.18 output per 1M tokens; provider documentation
- NanoGPT — model id inclusionai/ling-3.0-flash; $0.075 input / $0.22 output per 1M tokens; provider documentation