Voxtral Small 24B

by Mistral

Voxtral Small 24B is Mistral's first audio input model for instruct use cases, merging speech understanding with full language model capabilities. Built on Mistral Small 3.1, the 24B model preserves strong text performance while adding state-of-the-art audio understanding for transcription, translation, and speech-driven tasks. With a 32k token context window, Voxtral Small processes audio up to 30 minutes for transcription or 40 minutes for understanding and summarization. The model supports Q&A directly from audio, structured outputs, and function calling straight from voice input, enabling voice-first agentic workflows without separately chaining ASR and an LLM. Natively multilingual with automatic language detection, Voxtral Small delivers competitive performance across English, Spanish, French, Portuguese, Hindi, German, Dutch, and Italian. Released under Apache 2.0, the model is available on Hugging Face and via API.

Key info

Input
Output
Features
Context window
Max output

Available routes

No routes currently available — Voxtral Small 24B isn't routed through the Opper gateway right now. It may return.

Contact us about this model →

Available models from Mistral

Start building with 700+ models

One API key. Every major provider. Up and running in minutes.

Get startedView Documentation
Voxtral Small 24B by Mistral — not currently on Opper | Opper AI