Gemma 4 12B Instruct
Google's Gemma 4 12B Instruct is an open-weight multimodal model for text, image, audio, and video understanding, with tool calling and structured output support.
Context
131,072
Max output
32,768
Input / 1M
$0.05
Output / 1M
$0.25
Providers
2
01Profile
Profile
Released
2026-08-01
Last updated
2026-08-01
Knowledge cutoff
—
Weights
Open
02Capabilities
Capabilities
Input
text · image · video · audio
Output
text
Reasoning
No
Tool calling
Yes
Structured output
Yes
Attachments
Yes
Cache read
—
Cache write
—
03Pricing
Pricing across 2 providers
Per 1M tokens, USD. Sorted by input price. First-party row highlighted.
04History
Change history
6+ recent events since 2026-08-01
2026-09-16Contextnano-gptContext window262,144→ 131,072
2026-08-31Price cutnano-gptinput / 1M$0.06→ $0.05
2026-08-31Price cutnano-gptoutput / 1M$0.30→ $0.25
2026-08-03Addednano-gptoffering→ new
2026-07-15ContextpioneerMax output4,096→ 32,768
2026-06-29Addedpioneeroffering→ new