nemotron-lightning-3.5-30b-a3b
Nemotron-Lightning-3.5-30B-A3B is a 30B-parameter Mixture-of-Experts language model (3B active) from NVIDIA's Nemotron-H family, built on a hybrid Mamba-Transformer architecture for efficient long-context inference. Like other models in the family, it responds to queries by first generating a reasoning trace and then concluding with a final response, with reasoning behavior configurable through a flag in the chat template. It includes a multi-token prediction (MTP) speculative decoding head for low-latency serving.
上下文
262,144
最大输出
262,144
输入 / 1M
$0.05
输出 / 1M
$0.20
供应商
1
01概览
概览
发布日期
2026-08-15
最后更新
2026-08-15
知识截止
—
权重
闭源
02能力
能力
输入
text
输出
text
推理
是
工具调用
是
结构化输出
是
附件
否
缓存读取
—
缓存写入
—
03价格
1 家供应商的价格
每百万 token,美元。按输入价排序,一方行高亮。
04历史
变更历史
自 2026-08-15 以来的 5+ 条事件
2026-08-27涨价requesty输入 / 1M$0.04→ $0.05
2026-08-27涨价requesty输出 / 1M$0.18→ $0.20
2026-08-20降价requesty输入 / 1M$0.05→ $0.04
2026-08-20降价requesty输出 / 1M$0.20→ $0.18
2026-08-18新增requesty上架→ 新增