mercury familyActive
Mercury 2.5 Preview
Mercury 2.5 Preview is Inception's latest and most intelligent diffusion language model. Instead of generating tokens strictly one at a time, it produces and refines multiple tokens in parallel, reaching up to 1,107 tokens per second on standard GPUs. It delivers a 10+ point intelligence gain over Mercury 2, with tunable reasoning, parallel tool calls, schema-aligned JSON output, and a 260K context window. It is built for latency-sensitive production work such as search agents, voice pipelines, customer support, rapid coding iteration, and coding subagents.
Context
260,000
Max output
65,536
Input / 1M
$0.04
Output / 1M
$0.15
Providers
1
01Profile
Profile
Released
2026-09-01
Last updated
2026-09-01
Knowledge cutoff
—
Weights
Closed
02Capabilities
Capabilities
Input
text
Output
text
Reasoning
Yes
Tool calling
Yes
Structured output
Yes
Attachments
No
Cache read
—
Cache write
—
03Pricing
Pricing across 1 providers
Per 1M tokens, USD. Sorted by input price. First-party row highlighted.
04History
Change history
7+ recent events since 2026-09-01
2026-09-08Removedkiloofferingremoved→ new
2026-09-08Price upopenrouterinput / 1M$0.04→ $0.20
2026-09-08Price upopenrouteroutput / 1M$0.15→ $0.75
2026-09-08Removedopenrouterofferingremoved→ new
2026-09-01Addedkilooffering→ new
2026-09-01Addednano-gptoffering→ new
2026-09-01Addedopenrouteroffering→ new