Voxtral Small 24B 2507
Open audio-language model for speech transcription, audio understanding, and voice-driven tool use
Context
128,000
Max output
32,768
Input / 1M
$0.0023
Output / 1M
$0.0023
Providers
6
01Profile
Profile
Released
2025-07-15
Last updated
2025-07-15
Knowledge cutoff
—
Weights
Open
02Capabilities
Capabilities
Input
text · audio
Output
text
Reasoning
No
Tool calling
Yes
Structured output
Yes
Attachments
Yes
Cache read
—
Cache write
—
03Pricing
Pricing across 7 providers
Per 1M tokens, USD. Sorted by input price. First-party row highlighted.
04History
Change history
12+ recent events since 2025-07-15
2026-09-12Addededenaioffering→ new
2026-09-12Addededenaioffering→ new
2026-09-12ContextkiloContext window32,768→ new
2026-09-12ContextopenrouterContext window32,768→ new
2026-09-11Price cutamazon-bedrockinput / 1M$0.15→ $0.10
2026-09-11Price cutamazon-bedrockoutput / 1M$0.35→ $0.30
2026-09-11Contextamazon-bedrockContext window32,000→ new
2026-08-28ContextkiloContext window32,000→ 32,768
2026-08-28ContextkiloMax output25,600→ 26,214
2026-08-28ContextopenrouterContext window32,000→ 32,768
2026-08-28ContextopenrouterMax output25,600→ 26,214
2026-08-25ContextkiloMax output6,400→ 25,600