All AI models · LLM benchmarks · Methodology

Meta Llama 3.1 8B Instruct Turbo

Compact Llama instruction model for fast chat and local deployment

Model facts

Context window
128K tokens
Maximum output
128K tokens
Cheapest paid input
$0.02 per 1M tokens
Cheapest paid output
$0.03 per 1M tokens
Open weights
no
Providers
1

Capabilities and modalities

tool calling, temperature control, text.

Observed price history

  • 2026-08-24: $0.02 input / $0.03 output per 1M tokens

API providers

  • Helicone — model id llama-3.1-8b-instruct-turbo; $0.02 input / $0.03 output per 1M tokens; provider documentation