Active
Vision Medium
Balanced multimodal model pairing a 1M-token context window with deeper reasoning for analysis, content creation, and tool use across text, image, audio, video, and PDF inputs.
Context
1,000,000
Max output
65,536
Input / 1M
$4.21
Output / 1M
$12.63
Providers
1
01Profile
Profile
Released
2024-05-15
Last updated
2026-09
Knowledge cutoff
—
Weights
Closed
02Capabilities
Capabilities
Input
text · image · audio · video · pdf
Output
text
Reasoning
Yes
Tool calling
Yes
Structured output
Yes
Attachments
Yes
Cache read
—
Cache write
—
03Pricing
Pricing across 1 providers
Per 1M tokens, USD. Sorted by input price. First-party row highlighted.
04History
Change history
1+ recent events since 2024-05-15
2026-09-14Addedvisparkoffering→ new