nemotron-lightning-3.5-30b-a3b
Nemotron-Lightning-3.5-30B-A3B is a 30B-parameter Mixture-of-Experts language model (3B active) from NVIDIA's Nemotron-H family, built on a hybrid Mamba-Transformer architecture for efficient long-context inference. Like other models in the family, it responds to queries by first generating a reasoning trace and then concluding with a final response, with reasoning behavior configurable through a flag in the chat template. It includes a multi-token prediction (MTP) speculative decoding head for low-latency serving.
上下文
262,144
最大輸出
262,144
輸入 / 1M
$0.05
輸出 / 1M
$0.20
供應商
1
01概覽
概覽
釋出日期
2026-08-15
最後更新
2026-08-15
知識截止
—
權重
閉源
02能力
能力
輸入
text
輸出
text
推理
是
工具呼叫
是
結構化輸出
是
附件
否
快取讀取
—
快取寫入
—
03價格
1 家供應商的價格
每百萬 token,美元。按輸入價排序,一方行高亮。
04歷史
變更歷史
自 2026-08-15 以來的 5+ 條事件
2026-08-27漲價requesty輸入 / 1M$0.04→ $0.05
2026-08-27漲價requesty輸出 / 1M$0.18→ $0.20
2026-08-20降價requesty輸入 / 1M$0.05→ $0.04
2026-08-20降價requesty輸出 / 1M$0.20→ $0.18
2026-08-18新增requesty上架→ 新增