K2-Horizon-7B
IFM's open-weight 7B model for coding, reasoning, and chat. Heabsy serves an FP8 deployment on a dedicated RTX 5090 with speculative decoding in the United States, with a 131K context window and an 8K output limit. Zero data retention is not verified for this deployment.
Model facts
- Context window
- 131K tokens
- Maximum output
- 8K tokens
- Cheapest paid input
- $0.05 per 1M tokens
- Cheapest paid output
- $0.2 per 1M tokens
- Open weights
- yes
- Providers
- 1
Capabilities and modalities
reasoning, tool calling, open weights, text.
Local hardware estimate
At 4K context: Q4_K_M 6.5 GiB, Q8_0 10.7 GiB, F16 18.5 GiB working memory.
- Parameters
- 9.0 billion
- Architecture
- k2horizon
- Model license
- apache-2.0
Planning estimate adapted from the MIT-licensed whichllm estimator using metadata from Hugging Face; not a vendor minimum requirement.
Observed price history
- 2026-09-10: $0.05 input / $0.2 output per 1M tokens
API providers
- NanoGPT — model id k2-horizon-7b; $0.05 input / $0.2 output per 1M tokens; provider documentation