Voxtral Mini 3B 2507
Open audio-language model for speech transcription, audio understanding, and voice-driven tool use
Context
128,000
Max output
32,768
Input / 1M
$0.04
Output / 1M
$0.04
Providers
2
01Profile
Profile
Released
2025-07-15
Last updated
2025-07-15
Knowledge cutoff
—
Weights
Open
02Capabilities
Capabilities
Input
text · audio
Output
text
Reasoning
No
Tool calling
Yes
Structured output
Yes
Attachments
Yes
Cache read
—
Cache write
—
03Pricing
Pricing across 3 providers
Per 1M tokens, USD. Sorted by input price. First-party row highlighted.
04History
Change history
4+ recent events since 2025-07-15
2026-09-12Addededenaioffering→ new
2026-09-12Addededenaioffering→ new
2026-09-11Contextamazon-bedrockContext window128,000→ new
2025-12-13Addedamazon-bedrockoffering→ new