All AI models · LLM benchmarks · Methodology

InclusionAI

Ling 3.0 Flash

*Ling-3.0-flash* is a *124B-parameter Mixture-of-Experts (MoE) model*, with approximately *5.1B parameters activated per token*. The model is designed with *token efficiency and production-scale agentic inference* as key priorities, enabling developers...

Model facts

Context window
262K tokens
Maximum output
33K tokens
Cheapest paid input
$0.021 per 1M tokens
Cheapest paid output
$0.063 per 1M tokens
Open weights
yes
Providers
5

Capabilities and modalities

reasoning, tool calling, temperature control, open weights, text.

Local hardware estimate

At 4K context: Q4_K_M 74.5 GiB, Q8_0 133.9 GiB, F16 245.2 GiB working memory.

Parameters
127.5 billion
Architecture
bailingmoev3
Model license
mit

Planning estimate adapted from the MIT-licensed whichllm estimator using metadata from Hugging Face; not a vendor minimum requirement.

Artificial Analysis benchmarks

Compare this model on the full LLM benchmark leaderboard.

ConfigurationAAICodingMathOutput tokens/s
default37.850.6not measured336.0

Observed price history

  • 2026-08-24: $0.021 input / $0.063 output per 1M tokens

API providers

  • OpenRouter — model id inclusionai/ling-3.0-flash; $0.021 input / $0.063 output per 1M tokens; provider documentation
  • DevPass (LLM Gateway) (InclusionAI) — model id ling-3.0-flash; $0.06 input / $0.18 output per 1M tokens; provider documentation
  • Kilo Gateway — model id inclusionai/ling-3.0-flash; $0.06 input / $0.18 output per 1M tokens; provider documentation
  • Vercel AI Gateway — model id inclusionai/ling-3.0-flash; $0.06 input / $0.18 output per 1M tokens; provider documentation
  • NanoGPT — model id inclusionai/ling-3.0-flash; $0.075 input / $0.22 output per 1M tokens; provider documentation