Alibaba
Qwen3.8 2.4T A95B
Open-weight sparse MoE (2.4T total, 95B active), the open-weight twin of Qwen3.8 Max for coding, research, complex reasoning, and agentic workflows
Model facts
- Context window
- 262K tokens
- Maximum output
- 262K tokens
- Cheapest paid input
- $2 per 1M tokens
- Cheapest paid output
- $6 per 1M tokens
- Open weights
- yes
- Providers
- 15
Capabilities and modalities
reasoning, tool calling, structured output, attachments, temperature control, open weights, text, image, video, pdf.
Local hardware estimate
At 4K context: Q4_K_M 1412.6 GiB, Q8_0 2551.7 GiB, F16 4687.5 GiB working memory.
- Parameters
- 2446.2 billion
- Architecture
- qwen3_5moe
- Model license
- other
Planning estimate adapted from the MIT-licensed whichllm estimator using metadata from Hugging Face; not a vendor minimum requirement.
Artificial Analysis benchmarks
Compare this model on the full LLM benchmark leaderboard.
| Configuration | AAI | Coding | Math | Output tokens/s |
|---|---|---|---|---|
| default | 57.7 | 71.9 | not measured | 40.2 |
Observed price history
- 2026-08-24: $1.8 input / $5.4 output per 1M tokens
- 2026-08-28: $2 input / $6 output per 1M tokens
API providers
- AIHubMix — model id qwen3.8-2.4t-a95b; $2 input / $6 output per 1M tokens; provider documentation
- Charm Hyper — model id qwen3.8-2.4t-a95b; $2 input / $6 output per 1M tokens; provider documentation
- Deep Infra — model id Qwen/Qwen3.8-2.4T-A95B; $2 input / $6 output per 1M tokens; provider documentation
- Eden AI — model id qwen/qwen3.8-2.4t-a95b; $2 input / $6 output per 1M tokens; provider documentation
- Kilo Gateway — model id qwen/qwen3.8-2.4t-a95b; $2 input / $6 output per 1M tokens; provider documentation
- NanoGPT (Max) — model id qwen/qwen3.8-2.4t-a95b; $2 input / $6 output per 1M tokens; provider documentation
- OpenRouter — model id qwen/qwen3.8-2.4t-a95b; $2 input / $6 output per 1M tokens; provider documentation
- OpenRouter (batch) — model id qwen/qwen3.8-2.4t-a95b:batch; $2 input / $6 output per 1M tokens; provider documentation
- Requesty — model id qwen3.8-2.4T-A95B; $2 input / $6 output per 1M tokens; provider documentation
- RunInfra (NVFP4) — model id Inferact/Qwen3.8-2.4T-A95B-NVFP4; $2 input / $6 output per 1M tokens; provider documentation
- Vercel AI Gateway — model id alibaba/qwen3.8-2.4t-a95b; $2 input / $6 output per 1M tokens; provider documentation
- Cortecs — model id qwen3.8-2.4t-a95b; $2.5 input / $6 output per 1M tokens; provider documentation
- Requesty (EU) — model id qwen3.8-2.4T-A95B@eu; $2.5 input / $6 output per 1M tokens; provider documentation
- Hugging Face — model id Qwen/Qwen3.8-2.4T-A95B; $2.5 input / $6.25 output per 1M tokens; provider documentation
- Merge Gateway — model id qwen/qwen3.8-2.4t-a95b; $2.5 input / $6.25 output per 1M tokens; provider documentation