All AI models · LLM benchmarks · Methodology

voxtral-small-2507

Voxtral Small is a multimodal model with audio input, combining advanced speech capabilities with strong text performance for transcription, translation, and audio understanding.

Model facts

Context window
32K tokens
Maximum output
32K tokens
Cheapest paid input
$0.111 per 1M tokens
Cheapest paid output
$0.334 per 1M tokens
Open weights
no
Providers
1

Capabilities and modalities

tool calling, structured output, attachments, text, audio.

Observed price history

  • 2026-08-24: $0.111 input / $0.334 output per 1M tokens

API providers