All AI models · LLM benchmarks · Methodology

Meta

Llama 3.1 70B

Open Llama instruction model for multilingual chat, reasoning, and coding

Model facts

Context window
128K tokens
Maximum output
16K tokens
Cheapest paid input
$0.4 per 1M tokens
Cheapest paid output
$0.4 per 1M tokens
Open weights
yes
Providers
7

Capabilities and modalities

tool calling, structured output, temperature control, open weights, text.

Local hardware estimate

At 4K context: Q4_K_M 44.2 GiB, Q8_0 77.0 GiB, F16 138.6 GiB working memory.

Parameters
70.6 billion
Architecture
llama
Model license
llama3.1

Planning estimate adapted from the MIT-licensed whichllm estimator using metadata from Hugging Face; not a vendor minimum requirement.

Artificial Analysis benchmarks

Compare this model on the full LLM benchmark leaderboard.

ConfigurationAAICodingMathOutput tokens/s
default6.5not measured4.0not measured

Observed price history

  • 2026-08-24: $0.4 input / $0.4 output per 1M tokens

API providers

  • Nvidia — model id meta/llama-3.1-70b-instruct; free input / free output per 1M tokens; provider documentation
  • OpenRouter — model id meta-llama/llama-3.1-70b-instruct; $0.4 input / $0.4 output per 1M tokens; provider documentation
  • Kilo Gateway — model id meta-llama/llama-3.1-70b-instruct; $0.4 input / $0.4 output per 1M tokens; provider documentation
  • Amazon Bedrock — model id meta.llama3-1-70b-instruct-v1:0; $0.72 input / $0.72 output per 1M tokens; provider documentation
  • DevPass (LLM Gateway) — model id llama-3.1-70b-instruct; $0.72 input / $0.72 output per 1M tokens; provider documentation
  • Vercel AI Gateway — model id meta/llama-3.1-70b; $0.72 input / $0.72 output per 1M tokens; provider documentation
  • Weights & Biases — model id meta-llama/Llama-3.1-70B-Instruct; $0.8 input / $0.8 output per 1M tokens; provider documentation