GLM 4.1V Thinking Flash
Compact GPT model for low-latency assistance and high-volume workloads
Context
64,000
Max output
8,192
Input / 1M
$0.30
Output / 1M
$0.30
Providers
1
01Profile
Profile
Released
2025-07-09
Last updated
2025-07-09
Knowledge cutoff
—
Weights
Closed
02Capabilities
Capabilities
Input
text · image
Output
text
Reasoning
No
Tool calling
No
Structured output
No
Attachments
Yes
Cache read
—
Cache write
—
03Pricing
Pricing across 1 providers
Per 1M tokens, USD. Sorted by input price. First-party row highlighted.
04History
Change history
1+ recent events since 2025-07-09
2026-02-17Addednano-gptoffering→ new