All AI models · LLM benchmarks · Methodology

Meta

Llama 3.3 70B

Open Llama instruction model for multilingual chat, reasoning, and coding

Model facts

Context window
128K tokens
Maximum output
4K tokens
Cheapest paid input
$0.05 per 1M tokens
Cheapest paid output
$0.23 per 1M tokens
Open weights
yes
Providers
44

Capabilities and modalities

tool calling, structured output, attachments, temperature control, open weights, text.

Local hardware estimate

At 4K context: Q4_K_M 44.2 GiB, Q8_0 77.0 GiB, F16 138.6 GiB working memory.

Parameters
70.6 billion
Architecture
llama
Model license
llama3.3

Planning estimate adapted from the MIT-licensed whichllm estimator using metadata from Hugging Face; not a vendor minimum requirement.

Artificial Analysis benchmarks

Compare this model on the full LLM benchmark leaderboard.

ConfigurationAAICodingMathOutput tokens/s
default9.311.97.785.6

Observed price history

  • 2026-08-24: $0.05 input / $0.23 output per 1M tokens

API providers

  • Llama — model id llama-3.3-70b-instruct; free input / free output per 1M tokens; provider documentation
  • Nvidia — model id meta/llama-3.3-70b-instruct; free input / free output per 1M tokens; provider documentation
  • Pendra — model id llama3.3:70b; free input / free output per 1M tokens; provider documentation
  • Vercel AI Gateway — model id meta/llama-3.3-70b; free input / free output per 1M tokens; provider documentation
  • NanoGPT — model id meta-llama/llama-3.3-70b-instruct; $0.05 input / $0.23 output per 1M tokens; provider documentation
  • Meganova — model id meta-llama/Llama-3.3-70B-Instruct; $0.1 input / $0.3 output per 1M tokens; provider documentation
  • OpenRouter — model id meta-llama/llama-3.3-70b-instruct; $0.1 input / $0.32 output per 1M tokens; provider documentation
  • Eden AI — model id deepinfra/meta-llama/Llama-3.3-70B-Instruct; $0.1 input / $0.32 output per 1M tokens; provider documentation
  • Kilo Gateway — model id meta-llama/llama-3.3-70b-instruct; $0.1 input / $0.32 output per 1M tokens; provider documentation
  • IO.NET — model id meta-llama/Llama-3.3-70B-Instruct; $0.13 input / $0.38 output per 1M tokens; provider documentation
  • Helicone — model id llama-3.3-70b-instruct; $0.13 input / $0.39 output per 1M tokens; provider documentation
  • Cortecs — model id llama-3.3-70b-instruct; $0.129 input / $0.399 output per 1M tokens; provider documentation
  • Eden AI — model id nebius/meta-llama/Llama-3.3-70B-Instruct; $0.13 input / $0.4 output per 1M tokens; provider documentation
  • DevPass (LLM Gateway) — model id llama-3.3-70b-instruct; $0.135 input / $0.4 output per 1M tokens; provider documentation
  • NovitaAI — model id meta-llama/llama-3.3-70b-instruct; $0.135 input / $0.4 output per 1M tokens; provider documentation
  • Merge Gateway — model id meta/llama-3.3-70b-instruct; $0.22 input / $0.5 output per 1M tokens; provider documentation
  • Crusoe — model id meta-llama/Llama-3.3-70B-Instruct; $0.25 input / $0.75 output per 1M tokens; provider documentation
  • STACKIT — model id cortecs/Llama-3.3-70B-Instruct-FP8-Dynamic; $0.53 input / $0.76 output per 1M tokens; provider documentation
  • DigitalOcean — model id llama3.3-70b-instruct; $0.65 input / $0.65 output per 1M tokens; provider documentation
  • Abacus — model id meta-llama/Meta-Llama-3.3-70B-Instruct; $0.59 input / $0.79 output per 1M tokens; provider documentation
  • Groq — model id llama-3.3-70b-versatile; $0.59 input / $0.79 output per 1M tokens; provider documentation
  • Hugging Face — model id meta-llama/Llama-3.3-70B-Instruct; $0.59 input / $0.79 output per 1M tokens; provider documentation
  • Azure — model id llama-3.3-70b-instruct; $0.71 input / $0.71 output per 1M tokens; provider documentation
  • Azure Cognitive Services — model id llama-3.3-70b-instruct; $0.71 input / $0.71 output per 1M tokens; provider documentation
  • Weights & Biases — model id meta-llama/Llama-3.3-70B-Instruct; $0.71 input / $0.71 output per 1M tokens; provider documentation
  • Amazon Bedrock — model id meta.llama3-3-70b-instruct-v1:0; $0.72 input / $0.72 output per 1M tokens; provider documentation
  • Vertex — model id meta/llama-3.3-70b-instruct-maas; $0.72 input / $0.72 output per 1M tokens; provider documentation
  • OVHcloud AI Endpoints — model id meta-llama-3_3-70b-instruct; $0.74 input / $0.74 output per 1M tokens; provider documentation
  • watsonx.ai — model id meta-llama/llama-3-3-70b-instruct; $0.753 input / $0.753 output per 1M tokens; provider documentation
  • Eden AI — model id ionos/meta-llama/Llama-3.3-70B-Instruct; $0.755 input / $0.755 output per 1M tokens; provider documentation
  • Charm Hyper — model id llama-3.3-70b-instruct; $0.607 input / $1.04 output per 1M tokens; provider documentation
  • Pioneer — model id meta-llama/Llama-3.3-70B-Instruct; $0.9 input / $0.9 output per 1M tokens; provider documentation
  • Scaleway — model id llama-3.3-70b-instruct; $0.9 input / $0.9 output per 1M tokens; provider documentation
  • Neon — model id meta-llama-3-3-70b-instruct; $0.5 input / $1.5 output per 1M tokens; provider documentation
  • Together AI — model id meta-llama/Llama-3.3-70B-Instruct-Turbo; $1.04 input / $1.04 output per 1M tokens; provider documentation
  • Eden AI — model id scaleway/llama-3.3-70b-instruct; $1.05 input / $1.05 output per 1M tokens; provider documentation
  • evroc — model id nvidia/Llama-3.3-70B-Instruct-FP8; $1.15 input / $1.15 output per 1M tokens; provider documentation
  • GreenPT — model id llama-3.3-70b-instruct; $1.25 input / $1.25 output per 1M tokens; provider documentation
  • Regolo AI — model id llama-3.3-70b-instruct; $0.6 input / $2.7 output per 1M tokens; provider documentation
  • Venice AI — model id llama-3.3-70b; $0.7 input / $2.8 output per 1M tokens; provider documentation
  • NanoGPT — model id TEE/llama3-3-70b; $1.75 input / $2.75 output per 1M tokens; provider documentation
  • Tinfoil — model id llama3-3-70b; $1.75 input / $2.75 output per 1M tokens; provider documentation
  • CloudFerro Sherlock — model id meta-llama/Llama-3.3-70B-Instruct; $2.92 input / $2.92 output per 1M tokens; provider documentation
  • Snowflake Cortex — model id snowflake-llama3.3-70b; not published input / not published output per 1M tokens; provider documentation