All AI models · LLM benchmarks · Methodology

Meta

Llama 3.1 405B

Llama 3.1 405B hosted by IONOS in Berlin, Germany. Zero data retention.

Model facts

Context window
131K tokens
Maximum output
128K tokens
Cheapest paid input
$1.95 per 1M tokens
Cheapest paid output
$1.95 per 1M tokens
Open weights
yes
Providers
2

Capabilities and modalities

reasoning, tool calling, structured output, temperature control, open weights, text.

Artificial Analysis benchmarks

Compare this model on the full LLM benchmark leaderboard.

ConfigurationAAICodingMathOutput tokens/s
default7.3not measured3.0not measured

Observed price history

  • 2026-09-11: $1.95 input / $1.95 output per 1M tokens

API providers

  • Cortecs — model id llama-3.1-405b-instruct; $1.95 input / $1.95 output per 1M tokens; provider documentation
  • NanoGPT — model id meta-llama/llama-3.1-405b-instruct; $2.03 input / $2.03 output per 1M tokens; provider documentation