Voxtral Mini 2507
Voxtral Mini (July 2025) is a 3B open-weight speech-understanding model released under Apache 2.0 and built on the Ministral-3B language backbone. It combines transcription with native semantic understanding of audio, so a single model can both transcribe speech and answer questions, summarize, and extract insights from it without chaining a separate speech recognizer and language model. With a 32K token context window, Voxtral Mini handles audio up to 30 minutes for transcription or 40 minutes for understanding tasks, and is natively multilingual across widely used languages. It supports a dedicated transcription mode alongside built-in question answering and summarization. The model is sized for cost-efficient, responsive voice applications, making it a fit for voice agents, conversational AI, and agentic workflows that need both transcription accuracy and semantic comprehension of speech.
Key info
Available routes
No routes currently available — Voxtral Mini 2507 isn't routed through the Opper gateway right now. It may return.
Contact us about this model →