LLM List: AI models compared
Browse all LLMs in one comprehensive list: 1,781 large language models and AI models across 214 API providers. Compare token prices, context windows, capabilities, benchmark scores, open-weight availability, and local hardware requirements. Updated every six hours.
Explore the LLM benchmark leaderboard · How prices, benchmarks and model identities are calculated
- Gemma 4 31B IT TEE — 262K context; $0.15 input / $0.46 output per 1M tokens
- Gemma 4 31B MeroMero v2 — 262K context; $0.1 input / $0.45 output per 1M tokens
- Gemma 4 31B MeroMero v2 Thinking — 262K context; $0.1 input / $0.45 output per 1M tokens
- Gemma 4 31B Queen — 262K context; $0.306 input / $0.306 output per 1M tokens
- Gemma 4 31B Thinking by Google — 262K context; $0.1 input / $0.35 output per 1M tokens
- Gemma 4 31B Thinking TEE — 262K context; $0.4 input / $1 output per 1M tokens
- gemma 4 31B turbo TEE by Google — 131K context; $0.12 input / $0.37 output per 1M tokens
- Gemma 4 E2B Instruct by Google — 131K context; $0.02 input / $0.1 output per 1M tokens
- Gemma 4 E2B IT by Google — 33K context; $0.1 input / $0.1 output per 1M tokens
- Gemma 4 E4B Instruct by Google — 33K context; $0.04 input / $0.2 output per 1M tokens
- Gemma 4 E4B IT by Google — 131K context; $0.02 input / $0.1 output per 1M tokens
- Gemma 4 Uncensored — 256K context; $0.163 input / $0.5 output per 1M tokens
- Gemma Sea Lion V4 27B It by AI Singapore — 128K context; $0.351 input / $0.555 output per 1M tokens
- Gemsicle — 262K context; $0.1 input / $0.45 output per 1M tokens
- GLiGuard LLM Guardrails 300M by Fastino — 8K context; $0.15 input / $0.15 output per 1M tokens
- gliner-pii — 128K context; free input / free output per 1M tokens
- GLiNER2 Base by Fastino — 8K context; $0.15 input / $0.15 output per 1M tokens
- GLiNER2 Large by Fastino — 8K context; $0.15 input / $0.15 output per 1M tokens
- GLiNER2 Multi by Fastino — 8K context; $0.15 input / $0.15 output per 1M tokens
- GLiNER2 Multi Large by Fastino — 8K context; $0.15 input / $0.15 output per 1M tokens
- GLiNER2 Privacy Filter PII (Multi) by Fastino — 8K context; $0.15 input / $0.15 output per 1M tokens
- GLM 4 32B 0414 by THUDM — 128K context; $0.2 input / $0.2 output per 1M tokens
- GLM 4 9B 0414 by THUDM — 32K context; $0.2 input / $0.2 output per 1M tokens
- GLM 4 Air 0111 — 128K context; $0.139 input / $0.139 output per 1M tokens
- GLM 4 Plus 0111 — 128K context; $10 input / $10 output per 1M tokens
- GLM 4.1V Thinking Flash — 64K context; $0.3 input / $0.3 output per 1M tokens
- GLM 4.1V Thinking FlashX — 64K context; $0.3 input / $0.3 output per 1M tokens
- GLM 4.5 (Thinking) by Z.AI — 128K context; $0.3 input / $1.3 output per 1M tokens
- GLM 4.5 Air (Thinking) by Z.AI — 128K context; $0.12 input / $0.8 output per 1M tokens
- GLM 4.5 FP8 by Z.AI — 131K context; $0.2 input / $0.8 output per 1M tokens
- GLM 4.5V Thinking by Z.AI — 66K context; $0.6 input / $1.8 output per 1M tokens
- GLM 4.6 Derestricted v5 — 131K context; $0.4 input / $1.5 output per 1M tokens
- GLM 4.6 Original by Z.AI — 256K context; $0.35 input / $1.4 output per 1M tokens
- GLM 4.6 Thinking by Z.AI — 2K context; $0.35 input / $1.4 output per 1M tokens
- GLM 4.6 Turbo by Z.AI — 205K context; $1 input / $3 output per 1M tokens
- GLM 4.6 Turbo (Thinking) by Z.AI — 205K context; $1 input / $3 output per 1M tokens
- GLM 4.6V Flash by Z.AI — 2K context; $0.3 input / $0.9 output per 1M tokens
- GLM 4.6V Original by Z.AI — 128K context; $0.6 input / $0.9 output per 1M tokens
- GLM 4.7 Flash Heretic — 2K context; $0.07 input / $0.4 output per 1M tokens
- GLM 4.7 Flash Original by Z.AI — 2K context; $0.07 input / $0.4 output per 1M tokens
- GLM 4.7 Flash Original Thinking by Z.AI — 2K context; $0.07 input / $0.4 output per 1M tokens
- GLM 4.7 Flash Thinking by Z.AI — 2K context; $0.07 input / $0.4 output per 1M tokens
- GLM 4.7 Original by Z.AI — 2K context; $0.6 input / $2.2 output per 1M tokens
- GLM 4.7 Original Thinking by Z.AI — 2K context; $0.6 input / $2.2 output per 1M tokens
- GLM 4.7 Thinking by Z.AI — 2K context; $0.2 input / $0.8 output per 1M tokens
- GLM 5 Original by Z.AI — 2K context; $1 input / $3.2 output per 1M tokens
- GLM 5 Original Thinking by Z.AI — 2K context; $1 input / $3.2 output per 1M tokens
- GLM 5 Thinking by Z.AI — 2K context; $0.5 input / $2.55 output per 1M tokens
- GLM 5 Vision Turbo — 2K context; $0.704 input / $3.1 output per 1M tokens
- GLM 5.1 TEE by Z.AI — 203K context; $0.98 input / $3.08 output per 1M tokens
- GLM 5.1 Thinking by Z.AI — 2K context; $0.75 input / $2.6 output per 1M tokens
- GLM 5.1 Thinking TEE — 203K context; $1.5 input / $5.25 output per 1M tokens
- GLM 5.2 Fast by Z.AI — 1M context; $1.45 input / $4.5 output per 1M tokens
- GLM 5.2 Flex — 1.05M context; $0.943 input / $2.92 output per 1M tokens
- GLM 5.2 Nitro — 2K context; $0.8 input / $2.4 output per 1M tokens
- GLM 5.2 Short — 2K context; $1.45 input / $4.5 output per 1M tokens
- GLM 5.2 Short Fast — 2K context; $1.45 input / $4.5 output per 1M tokens
- GLM 5.2 Short Fast Flex — 2K context; $0.943 input / $2.92 output per 1M tokens
- GLM 5.2 Short Flex — 2K context; $0.943 input / $2.92 output per 1M tokens
- GLM 5.2 TEE by Z.AI — 1.05M context; $1.25 input / $3.95 output per 1M tokens
- GLM 5.2 Thinking by Z.AI — 1.05M context; $0.42 input / $1.32 output per 1M tokens
- GLM 5.2 Thinking TEE — 1.05M context; $1.4 input / $4.6 output per 1M tokens
- GLM 5.3 Fast by Z.AI — 1.05M context; $2.1 input / $6.6 output per 1M tokens
- GLM 5.3 Flash TEE — 1.05M context; $0.15 input / $0.5 output per 1M tokens
- GLM 5.3 Flash Uncensored by Z.AI — 1.05M context; $0.35 input / $1.4 output per 1M tokens
- GLM 5.3 FP4 — 1.05M context; $1.12 input / $4.4 output per 1M tokens
- GLM 5.3 TEE — 1.05M context; $1.4 input / $4.4 output per 1M tokens
- GLM 5.3 Thinking by Z.AI — 1.05M context; $1 input / $3.2 output per 1M tokens
- GLM 5V Turbo Thinking by Z.AI — 203K context; $1.2 input / $4 output per 1M tokens
- GLM Flash Latest by Z.AI — 1.31M context; $0.075 input / $0.25 output per 1M tokens
- GLM Latest by Z.AI — 1.31M context; $1 input / $3.2 output per 1M tokens
- GLM Z1 9B 0414 by THUDM — 32K context; $0.2 input / $0.2 output per 1M tokens
- GLM Z1 AirX — 32K context; $0.7 input / $0.7 output per 1M tokens
- GLM-4 32B (0414-128k) — 128K context; $0.1 input / $0.1 output per 1M tokens
- GLM-4 32B (0414-128k) (Z AI) by Z.AI — 128K context; $0.1 input / $0.1 output per 1M tokens
- GLM-4 Long — 1M context; $0.201 input / $0.201 output per 1M tokens
- GLM-4.5 by Z.AI — 131K context; $0.286 input / $1.14 output per 1M tokens
- GLM-4.5 AirX by Z.AI — 128K context; $0.572 input / $1.71 output per 1M tokens
- GLM-4.5 X by Z.AI — 128K context; $1.14 input / $2.29 output per 1M tokens
- GLM-4.5-Air by Z.AI — 131K context; $0.114 input / $0.286 output per 1M tokens
- GLM-4.5-Flash — 131K context; free input / free output per 1M tokens
- GLM-4.5V by Z.AI — 66K context; $0.29 input / $0.86 output per 1M tokens
- GLM-4.6 by Z.AI — 205K context; $0.286 input / $1.14 output per 1M tokens
- GLM-4.6V by Z.AI — 131K context; $0.14 input / $0.42 output per 1M tokens
- GLM-4.6V FlashX by Z.AI — 128K context; $0.02 input / $0.21 output per 1M tokens
- GLM-4.7 by Z.AI — 205K context; $0.2 input / $0.8 output per 1M tokens
- GLM-4.7 Free — 205K context; free input / free output per 1M tokens
- GLM-4.7-Flash by Z.AI — 2K context; $0.06 input / $0.4 output per 1M tokens
- GLM-4.7-FlashX by Z.AI — 2K context; $0.06 input / $0.4 output per 1M tokens
- glm-4.7-n — 205K context; not published input / not published output per 1M tokens
- GLM-5 by Z.AI — 203K context; $0.6 input / $1.92 output per 1M tokens
- GLM-5 Free — 205K context; free input / free output per 1M tokens
- GLM-5-Turbo by Z.AI — 2K context; $0.72 input / $3.2 output per 1M tokens
- GLM-5.1 by Z.AI — 2K context; $0.45 input / $2.15 output per 1M tokens
- GLM-5.1 FP8 by Z.AI — 203K context; $0.85 input / $3.3 output per 1M tokens
- GLM-5.2 by Z.AI — 1M context; $0.3 input / $1.05 output per 1M tokens
- GLM-5.2 Caveman — 1M context; $1.25 input / $5.02 output per 1M tokens
- GLM-5.2 Caveman Lite — 1M context; $1.25 input / $5.02 output per 1M tokens
- GLM-5.2 Caveman Ultra — 1M context; $1.25 input / $5.02 output per 1M tokens
- GLM-5.2 Highspeed — 1M context; free input / free output per 1M tokens