All AI models · LLM benchmarks · Methodology

Meta

Llama 4 Maverick 17B 128E Instruct FP8

Open multimodal Llama model for strong reasoning and fast responses

Model facts

Context window
1M tokens
Maximum output
16K tokens
Cheapest paid input
$0.25 per 1M tokens
Cheapest paid output
$1 per 1M tokens
Open weights
yes
Providers
5

Capabilities and modalities

tool calling, attachments, temperature control, open weights, text, image.

Observed price history

  • 2026-08-24: $0.25 input / $1 output per 1M tokens

API providers

  • Llama — model id llama-4-maverick-17b-128e-instruct-fp8; free input / free output per 1M tokens; provider documentation
  • Vercel AI Gateway — model id meta/llama-4-maverick; free input / free output per 1M tokens; provider documentation
  • Azure — model id llama-4-maverick-17b-128e-instruct-fp8; $0.25 input / $1 output per 1M tokens; provider documentation
  • Azure Cognitive Services — model id llama-4-maverick-17b-128e-instruct-fp8; $0.25 input / $1 output per 1M tokens; provider documentation
  • watsonx.ai — model id meta-llama/llama-4-maverick-17b-128e-instruct-fp8; $0.371 input / $1.48 output per 1M tokens; provider documentation