活跃
Vision Small
Fast, low-cost multimodal model for understanding text, images, audio, video, and PDFs, with tool calling and a 1M-token context window.
上下文
1,000,000
最大输出
65,536
输入 / 1M
$1.05
输出 / 1M
$3.16
供应商
1
01概览
概览
发布日期
2024-05-15
最后更新
2026-09
知识截止
—
权重
闭源
02能力
能力
输入
text · image · audio · video · pdf
输出
text
推理
是
工具调用
是
结构化输出
是
附件
是
缓存读取
—
缓存写入
—
03价格
1 家供应商的价格
每百万 token,美元。按输入价排序,一方行高亮。
04历史
变更历史
自 2024-05-15 以来的 1+ 条事件
2026-09-14新增vispark上架→ 新增