Voxtral Mini 2507

by Mistral

Voxtral Mini (July 2025) is a 3B open-weight speech-understanding model released under Apache 2.0 and built on the Ministral-3B language backbone. It combines transcription with native semantic understanding of audio, so a single model can both transcribe speech and answer questions, summarize, and extract insights from it without chaining a separate speech recognizer and language model. With a 32K token context window, Voxtral Mini handles audio up to 30 minutes for transcription or 40 minutes for understanding tasks, and is natively multilingual across widely used languages. It supports a dedicated transcription mode alongside built-in question answering and summarization. The model is sized for cost-efficient, responsive voice applications, making it a fit for voice agents, conversational AI, and agentic workflows that need both transcription accuracy and semantic comprehension of speech.

Key info

Input
Output
Features
Context window
33K
Max output
—

Available routes

No routes currently available — Voxtral Mini 2507 isn't routed through the Opper gateway right now. It may return.

Contact us about this model →

Available models from Mistral

Start building with 700+ models

One API key. Every major provider. Up and running in minutes.

Get startedView Documentation