InclusionAI
Ling 3.0 Flash VL
Ling 3.0 Flash VL is inclusionAI's native multimodal Mixture-of-Experts model with 124B total parameters and 5.5B active parameters per token. It combines image and video understanding with reasoning and tool use for document analysis, charts, visual verification, and interface-based agent tasks. Thinking is enabled by default and can be turned off in settings.
Model facts
- Context window
- 262K tokens
- Maximum output
- 33K tokens
- Cheapest paid input
- $0.06 per 1M tokens
- Cheapest paid output
- $0.18 per 1M tokens
- Open weights
- yes
- Providers
- 1
Capabilities and modalities
reasoning, tool calling, attachments, open weights, text, image, video.
Local hardware estimate
At 4K context: Q4_K_M 77.4 GiB, Q8_0 135.5 GiB, F16 244.5 GiB working memory.
- Parameters
- 124.8 billion
- Architecture
- bailingmoev3vl
- Model license
- mit
Planning estimate adapted from the MIT-licensed whichllm estimator using metadata from Hugging Face; not a vendor minimum requirement.
Observed price history
- 2026-09-09: $0.06 input / $0.18 output per 1M tokens
API providers
- NanoGPT — model id inclusionai/ling-3.0-flash-vl; $0.06 input / $0.18 output per 1M tokens; provider documentation