nemotron-lightning-3.5-30b-a3b
Nemotron-Lightning-3.5-30B-A3B is a 30B-parameter Mixture-of-Experts language model (3B active) from NVIDIA's Nemotron-H family, built on a hybrid Mamba-Transformer architecture for efficient long-context inference. Like other models in the family, it responds to queries by first generating a reasoning trace and then concluding with a final response, with reasoning behavior configurable through a flag in the chat template. It includes a multi-token prediction (MTP) speculative decoding head for low-latency serving.
Context
262,144
Max output
262,144
Input / 1M
$0.05
Output / 1M
$0.20
Providers
1
01Profile
Profile
Released
2026-08-15
Last updated
2026-08-15
Knowledge cutoff
—
Weights
Closed
02Capabilities
Capabilities
Input
text
Output
text
Reasoning
Yes
Tool calling
Yes
Structured output
Yes
Attachments
No
Cache read
—
Cache write
—
03Pricing
Pricing across 1 providers
Per 1M tokens, USD. Sorted by input price. First-party row highlighted.
04History
Change history
5+ recent events since 2026-08-15
2026-08-27Price uprequestyinput / 1M$0.04→ $0.05
2026-08-27Price uprequestyoutput / 1M$0.18→ $0.20
2026-08-20Price cutrequestyinput / 1M$0.05→ $0.04
2026-08-20Price cutrequestyoutput / 1M$0.20→ $0.18
2026-08-18Addedrequestyoffering→ new