All AI models · LLM benchmarks · Methodology

Meta

Llama 4 Maverick 17B FP8

Open multimodal Llama model for strong reasoning and fast responses

Model facts

Context window
1.05M tokens
Maximum output
16K tokens
Cheapest paid input
$0.2 per 1M tokens
Cheapest paid output
$0.8 per 1M tokens
Open weights
yes
Providers
1

Capabilities and modalities

structured output, attachments, open weights, text, image.

Observed price history

  • 2026-08-24: $0.2 input / $0.8 output per 1M tokens

API providers

  • Deep Infra — model id meta-llama/Llama-4-Maverick-17B-128E-Instruct-FP8; $0.2 input / $0.8 output per 1M tokens; provider documentation