Voxtral Small 2507

by Mistral

Voxtral Small (latest) aliases the newest Voxtral Small 24B snapshot, giving access to the most recent improvements in audio-native instruction following and agentic reasoning. It combines a 24B model for audio understanding with full language model capabilities for speech-driven applications. Built for multi-turn conversations that blend audio and text inputs, the model supports Q&A directly from audio, structured outputs, and function calling from spoken commands. The 32k token context window handles audio up to 30 minutes for transcription or 40 minutes for understanding. As an open-weight model under Apache 2.0, Voxtral Small works natively across eight major languages with automatic detection and is available both as a Hugging Face download and via API.

Key info

Input
Output
Features
Context window
33K
Max output
—

Available routes

No routes currently available — Voxtral Small 2507 isn't routed through the Opper gateway right now. It may return.

Contact us about this model →

Available models from Mistral

Start building with 700+ models

One API key for every major provider, up and running in minutes.

Get startedView Documentation