NVIDIAnemotron familyActive

nemotron-lightning-3.5-30b-a3b

Nemotron-Lightning-3.5-30B-A3B is a 30B-parameter Mixture-of-Experts language model (3B active) from NVIDIA's Nemotron-H family, built on a hybrid Mamba-Transformer architecture for efficient long-context inference. Like other models in the family, it responds to queries by first generating a reasoning trace and then concluding with a final response, with reasoning behavior configurable through a flag in the chat template. It includes a multi-token prediction (MTP) speculative decoding head for low-latency serving.

JSONEmbed chartCompare
Context
262,144
Max output
262,144
Input / 1M
$0.05
Output / 1M
$0.20
Providers
1
01Profile

Profile

Released
2026-08-15
Last updated
2026-08-15
Knowledge cutoff
Weights
Closed
02Capabilities

Capabilities

Input
text
Output
text
Reasoning
Yes
Tool calling
Yes
Structured output
Yes
Attachments
No
Cache read
Cache write
03Pricing

Pricing across 1 providers

Per 1M tokens, USD. Sorted by input price. First-party row highlighted.

ProviderProvider model idInputOutputCache readContext
requestynemotron-lightning-3.5-30b-a3b$0.05$0.20$0.01262,144
04History

Change history

5+ recent events since 2026-08-15

2026-08-27Price uprequestyinput / 1M$0.04→ $0.05
2026-08-27Price uprequestyoutput / 1M$0.18→ $0.20
2026-08-20Price cutrequestyinput / 1M$0.05→ $0.04
2026-08-20Price cutrequestyoutput / 1M$0.20→ $0.18
2026-08-18Addedrequestyoffering→ new
as of 2026-09-16
About ·Corrections ·Contact ·Privacy ·EN / 中文 / 繁體
HomeModelsLabsProvidersToolsRankingsChangesNews