Meta
Llama 3.3 70B
Open Llama instruction model for multilingual chat, reasoning, and coding
Model facts
- Context window
- 128K tokens
- Maximum output
- 4K tokens
- Cheapest paid input
- $0.05 per 1M tokens
- Cheapest paid output
- $0.23 per 1M tokens
- Open weights
- yes
- Providers
- 44
Capabilities and modalities
tool calling, structured output, attachments, temperature control, open weights, text.
Local hardware estimate
At 4K context: Q4_K_M 44.2 GiB, Q8_0 77.0 GiB, F16 138.6 GiB working memory.
- Parameters
- 70.6 billion
- Architecture
- llama
- Model license
- llama3.3
Planning estimate adapted from the MIT-licensed whichllm estimator using metadata from Hugging Face; not a vendor minimum requirement.
Artificial Analysis benchmarks
Compare this model on the full LLM benchmark leaderboard.
| Configuration | AAI | Coding | Math | Output tokens/s |
|---|---|---|---|---|
| default | 9.3 | 11.9 | 7.7 | 85.6 |
Observed price history
- 2026-08-24: $0.05 input / $0.23 output per 1M tokens
API providers
- Llama — model id llama-3.3-70b-instruct; free input / free output per 1M tokens; provider documentation
- Nvidia — model id meta/llama-3.3-70b-instruct; free input / free output per 1M tokens; provider documentation
- Pendra — model id llama3.3:70b; free input / free output per 1M tokens; provider documentation
- Vercel AI Gateway — model id meta/llama-3.3-70b; free input / free output per 1M tokens; provider documentation
- NanoGPT — model id meta-llama/llama-3.3-70b-instruct; $0.05 input / $0.23 output per 1M tokens; provider documentation
- Meganova — model id meta-llama/Llama-3.3-70B-Instruct; $0.1 input / $0.3 output per 1M tokens; provider documentation
- OpenRouter — model id meta-llama/llama-3.3-70b-instruct; $0.1 input / $0.32 output per 1M tokens; provider documentation
- Eden AI — model id deepinfra/meta-llama/Llama-3.3-70B-Instruct; $0.1 input / $0.32 output per 1M tokens; provider documentation
- Kilo Gateway — model id meta-llama/llama-3.3-70b-instruct; $0.1 input / $0.32 output per 1M tokens; provider documentation
- IO.NET — model id meta-llama/Llama-3.3-70B-Instruct; $0.13 input / $0.38 output per 1M tokens; provider documentation
- Helicone — model id llama-3.3-70b-instruct; $0.13 input / $0.39 output per 1M tokens; provider documentation
- Cortecs — model id llama-3.3-70b-instruct; $0.129 input / $0.399 output per 1M tokens; provider documentation
- Eden AI — model id nebius/meta-llama/Llama-3.3-70B-Instruct; $0.13 input / $0.4 output per 1M tokens; provider documentation
- DevPass (LLM Gateway) — model id llama-3.3-70b-instruct; $0.135 input / $0.4 output per 1M tokens; provider documentation
- NovitaAI — model id meta-llama/llama-3.3-70b-instruct; $0.135 input / $0.4 output per 1M tokens; provider documentation
- Merge Gateway — model id meta/llama-3.3-70b-instruct; $0.22 input / $0.5 output per 1M tokens; provider documentation
- Crusoe — model id meta-llama/Llama-3.3-70B-Instruct; $0.25 input / $0.75 output per 1M tokens; provider documentation
- STACKIT — model id cortecs/Llama-3.3-70B-Instruct-FP8-Dynamic; $0.53 input / $0.76 output per 1M tokens; provider documentation
- DigitalOcean — model id llama3.3-70b-instruct; $0.65 input / $0.65 output per 1M tokens; provider documentation
- Abacus — model id meta-llama/Meta-Llama-3.3-70B-Instruct; $0.59 input / $0.79 output per 1M tokens; provider documentation
- Groq — model id llama-3.3-70b-versatile; $0.59 input / $0.79 output per 1M tokens; provider documentation
- Hugging Face — model id meta-llama/Llama-3.3-70B-Instruct; $0.59 input / $0.79 output per 1M tokens; provider documentation
- Azure — model id llama-3.3-70b-instruct; $0.71 input / $0.71 output per 1M tokens; provider documentation
- Azure Cognitive Services — model id llama-3.3-70b-instruct; $0.71 input / $0.71 output per 1M tokens; provider documentation
- Weights & Biases — model id meta-llama/Llama-3.3-70B-Instruct; $0.71 input / $0.71 output per 1M tokens; provider documentation
- Amazon Bedrock — model id meta.llama3-3-70b-instruct-v1:0; $0.72 input / $0.72 output per 1M tokens; provider documentation
- Vertex — model id meta/llama-3.3-70b-instruct-maas; $0.72 input / $0.72 output per 1M tokens; provider documentation
- OVHcloud AI Endpoints — model id meta-llama-3_3-70b-instruct; $0.74 input / $0.74 output per 1M tokens; provider documentation
- watsonx.ai — model id meta-llama/llama-3-3-70b-instruct; $0.753 input / $0.753 output per 1M tokens; provider documentation
- Eden AI — model id ionos/meta-llama/Llama-3.3-70B-Instruct; $0.755 input / $0.755 output per 1M tokens; provider documentation
- Charm Hyper — model id llama-3.3-70b-instruct; $0.607 input / $1.04 output per 1M tokens; provider documentation
- Pioneer — model id meta-llama/Llama-3.3-70B-Instruct; $0.9 input / $0.9 output per 1M tokens; provider documentation
- Scaleway — model id llama-3.3-70b-instruct; $0.9 input / $0.9 output per 1M tokens; provider documentation
- Neon — model id meta-llama-3-3-70b-instruct; $0.5 input / $1.5 output per 1M tokens; provider documentation
- Together AI — model id meta-llama/Llama-3.3-70B-Instruct-Turbo; $1.04 input / $1.04 output per 1M tokens; provider documentation
- Eden AI — model id scaleway/llama-3.3-70b-instruct; $1.05 input / $1.05 output per 1M tokens; provider documentation
- evroc — model id nvidia/Llama-3.3-70B-Instruct-FP8; $1.15 input / $1.15 output per 1M tokens; provider documentation
- GreenPT — model id llama-3.3-70b-instruct; $1.25 input / $1.25 output per 1M tokens; provider documentation
- Regolo AI — model id llama-3.3-70b-instruct; $0.6 input / $2.7 output per 1M tokens; provider documentation
- Venice AI — model id llama-3.3-70b; $0.7 input / $2.8 output per 1M tokens; provider documentation
- NanoGPT — model id TEE/llama3-3-70b; $1.75 input / $2.75 output per 1M tokens; provider documentation
- Tinfoil — model id llama3-3-70b; $1.75 input / $2.75 output per 1M tokens; provider documentation
- CloudFerro Sherlock — model id meta-llama/Llama-3.3-70B-Instruct; $2.92 input / $2.92 output per 1M tokens; provider documentation
- Snowflake Cortex — model id snowflake-llama3.3-70b; not published input / not published output per 1M tokens; provider documentation