NVIDIAnemotron familyActive

nemotron-3-ultra

Largest Nemotron 3 model for maximum open-weight reasoning and agent accuracy

JSONEmbed chartCompare
Context
1,000,000
Max output
131,072
Input / 1M
$0.10
Output / 1M
$0.10
Providers
4
01Profile

Profile

Released
2026-06-04
Last updated
2026-06-23
Knowledge cutoff
Weights
Open
02Capabilities

Capabilities

Input
text
Output
text
Reasoning
Yes
Tool calling
Yes
Structured output
No
Attachments
No
Cache read
Cache write
03Pricing

Pricing across 3 providers

Per 1M tokens, USD. Sorted by input price. First-party row highlighted.

ProviderProvider model idInputOutputCache readContext
ollama-cloudnemotron-3-ultra$0.10$3.00$0.10262,144
routing-runnemotron-3-ultra$0.10$0.10Not listed131,072
requestynvidia-nemotron-3-ultraorg:nvidia$0.50$2.50Not listed262,144
opencodenemotron-3-ultra-freefreeFreeFreeFree1,000,000
04History

Change history

10+ recent events since 2026-06-04

2026-09-14Price upollama-cloudinput / 1M→ $0.10
2026-09-14Price upollama-cloudoutput / 1M→ $3.00
2026-08-27Price uprequestyinput / 1M$0.45→ $0.50
2026-08-27Price uprequestyoutput / 1M$2.25→ $2.50
2026-08-20Price cutrequestyinput / 1M$0.50→ $0.45
2026-08-20Price cutrequestyoutput / 1M$2.50→ $2.25
2026-08-18Addedrequestyoffering→ new
2026-07-09Addedrouting-runoffering→ new
2026-06-08Addedollama-cloudoffering→ new
2026-06-04Addedopencodeoffering→ new
as of 2026-09-16
About ·Corrections ·Contact ·Privacy ·EN / 中文 / 繁體
HomeModelsLabsProvidersToolsRankingsChangesNews