Vision Small
Fast, low-cost multimodal model for understanding text, images, audio, video, and PDFs, with tool calling and a 1M-token context window.
Model facts
- Context window
- 1M tokens
- Maximum output
- 66K tokens
- Cheapest paid input
- $1.05 per 1M tokens
- Cheapest paid output
- $3.16 per 1M tokens
- Open weights
- no
- Providers
- 1
Capabilities and modalities
reasoning, tool calling, structured output, attachments, temperature control, text, image, audio, video, pdf.
Observed price history
- 2026-09-15: $1.05 input / $3.16 output per 1M tokens
API providers
- Vispark — model id vispark/vision-small; $1.05 input / $3.16 output per 1M tokens; provider documentation