Voxtral Small 24B
Voxtral Small 24B is Mistral's first audio input model for instruct use cases, merging speech understanding with full language model capabilities. Built on Mistral Small 3.1, the 24B model preserves strong text performance while adding state-of-the-art audio understanding for transcription, translation, and speech-driven tasks. With a 32k token context window, Voxtral Small processes audio up to 30 minutes for transcription or 40 minutes for understanding and summarization. The model supports Q&A directly from audio, structured outputs, and function calling straight from voice input, enabling voice-first agentic workflows without separately chaining ASR and an LLM. Natively multilingual with automatic language detection, Voxtral Small delivers competitive performance across English, Spanish, French, Portuguese, Hindi, German, Dutch, and Italian. Released under Apache 2.0, the model is available on Hugging Face and via API.
Key info
Available routes
No routes currently available — Voxtral Small 24B isn't routed through the Opper gateway right now. It may return.
Contact us about this model →