GLM-4.5-Flash
Efficient GLM model for fast reasoning, coding, and agent workflows
Context
200,000
Max output
98,304
Input / 1M
Not listed
Output / 1M
Not listed
Providers
4
01Profile
Profile
Released
2025-07-28
Last updated
2025-07-28
Knowledge cutoff
2025-04
Weights
Open
02Capabilities
Capabilities
Input
text
Output
text
Reasoning
Yes
Tool calling
Yes
Structured output
No
Attachments
No
Cache read
—
Cache write
—
03Pricing
Pricing across 0 providers
Per 1M tokens, USD. Sorted by input price. First-party row highlighted.
04History
Change history
12+ recent events since 2025-07-28
2026-07-10Addedempiriolabsoffering→ new
2026-07-10Addedunorouteroffering→ new
2026-06-29Removedllmgatewayofferingremoved→ new
2026-06-24ContextllmgatewayContext window—→ 128,000
2026-06-03Price upllmgatewayinput / 1M—→ Not listed
2026-06-03Price upllmgatewayoutput / 1M—→ Not listed
2026-04-25Removedzai-coding-planofferingremoved→ new
2026-04-25Removedzhipuai-coding-planofferingremoved→ new
2026-04-18Price upllmgatewayinput / 1MNot listed→ new
2026-04-18Price upllmgatewayoutput / 1MNot listed→ new
2026-04-18ContextllmgatewayContext window128,000→ new
2026-04-18ContextllmgatewayMax output16,384→ new