Meta
Llama 4 Maverick 17B 128E Instruct FP8
Open multimodal Llama model for strong reasoning and fast responses
Model facts
- Context window
- 1M tokens
- Maximum output
- 16K tokens
- Cheapest paid input
- $0.25 per 1M tokens
- Cheapest paid output
- $1 per 1M tokens
- Open weights
- yes
- Providers
- 5
Capabilities and modalities
tool calling, attachments, temperature control, open weights, text, image.
Observed price history
- 2026-08-24: $0.25 input / $1 output per 1M tokens
API providers
- Llama — model id llama-4-maverick-17b-128e-instruct-fp8; free input / free output per 1M tokens; provider documentation
- Vercel AI Gateway — model id meta/llama-4-maverick; free input / free output per 1M tokens; provider documentation
- Azure — model id llama-4-maverick-17b-128e-instruct-fp8; $0.25 input / $1 output per 1M tokens; provider documentation
- Azure Cognitive Services — model id llama-4-maverick-17b-128e-instruct-fp8; $0.25 input / $1 output per 1M tokens; provider documentation
- watsonx.ai — model id meta-llama/llama-4-maverick-17b-128e-instruct-fp8; $0.371 input / $1.48 output per 1M tokens; provider documentation