Voxtral Small 2507
Voxtral Small (latest) aliases the newest Voxtral Small 24B snapshot, giving access to the most recent improvements in audio-native instruction following and agentic reasoning. It combines a 24B model for audio understanding with full language model capabilities for speech-driven applications. Built for multi-turn conversations that blend audio and text inputs, the model supports Q&A directly from audio, structured outputs, and function calling from spoken commands. The 32k token context window handles audio up to 30 minutes for transcription or 40 minutes for understanding. As an open-weight model under Apache 2.0, Voxtral Small works natively across eight major languages with automatic detection and is available both as a Hugging Face download and via API.
Key info
Available routes
No routes currently available — Voxtral Small 2507 isn't routed through the Opper gateway right now. It may return.
Contact us about this model →