MistralActive
pixtral-12b-2409
Pixtral 2409 12B is a state-of-the-art multimodal model with 12B parameters and a 400M vision encoder, natively trained on interleaved text and image data. It excels in tasks spanning vision-language reasoning, instruction following, and pure text understanding, making it highly effective for real-world multimodal applications.
Context
128,000
Max output
4,096
Input / 1M
$0.15
Output / 1M
$0.15
Providers
4
01Profile
Profile
Released
2024-11-09
Last updated
2026-03-17
Knowledge cutoff
—
Weights
Closed
02Capabilities
Capabilities
Input
text · image
Output
text
Reasoning
Yes
Tool calling
Yes
Structured output
Yes
Attachments
Yes
Cache read
—
Cache write
—
03Pricing
Pricing across 4 providers
Per 1M tokens, USD. Sorted by input price. First-party row highlighted.
04History
Change history
6+ recent events since 2024-11-09
2026-09-02ContextcortecsMax output128,000→ 4,096
2026-08-05Addedcortecsoffering→ new
2026-08-05Addedpioneeroffering→ new
2026-07-29Addedgreenptoffering→ new
2026-06-04ContextscalewayContext window128,000→ new
2025-10-20Addedscalewayoffering→ new