All AI models · LLM benchmarks · Methodology

Nous Research

Hermes 4 70B

Hermes 4 70B is a hybrid reasoning model from Nous Research, built on Meta-Llama-3.1-70B. It introduces the same hybrid mode as the larger 405B release, allowing the model to either...

Model facts

Context window
131K tokens
Maximum output
128K tokens
Cheapest paid input
$0.129 per 1M tokens
Cheapest paid output
$0.399 per 1M tokens
Open weights
yes
Providers
2

Capabilities and modalities

reasoning, tool calling, structured output, temperature control, open weights, text.

Local hardware estimate

At 4K context: Q4_K_M 44.2 GiB, Q8_0 77.0 GiB, F16 138.6 GiB working memory.

Parameters
70.6 billion
Architecture
llama
Model license
llama3

Planning estimate adapted from the MIT-licensed whichllm estimator using metadata from Hugging Face; not a vendor minimum requirement.

Observed price history

  • 2026-08-24: $0.129 input / $0.399 output per 1M tokens

API providers