Runtime is almost here: join TypeSafe AI, Cognition, DoorDash and more in SF. Limited seats left. Register now

Nemotron 3.5 Lightning 30B A3B NVFP4

Nemotron 3.5 Lightning is NVIDIA's open 30B Mixture-of-Experts model with 3B active parameters, built for high-volume, low-latency task execution in long-running agents. This NVFP4 checkpoint takes text input and offers a 1M-token context, configurable thinking, and OpenMDW-1.1 licensing.

Model

Nemotron 3.5 Lightning is a hybrid Mixture-of-Experts model that interleaves Mamba-2 and MoE layers with a handful of attention layers: 52 layers in total, 128 routed experts plus one shared expert with six experts active per token, and 30 billion total parameters of which 3 billion are active. It was pre-trained on over 20 trillion tokens with an NVFP4 recipe and supports context windows up to 1 million tokens. Multi-token prediction is baked in during pre-training. Thinking is on by default and can be disabled through chat-template kwargs, and the model supports tool calling and six languages: English, Spanish, French, German, Italian, and Japanese.

This is NVIDIA's own NVFP4 quantization, produced with Model Optimizer using a Four Over Six NVFP4 recipe: W4A16 on the routed and shared experts, with FP8 per-tensor dynamic scales on the Mamba projections and KV cache. It runs natively in FP4 on Blackwell GPUs and through W4A16 kernels on Hopper and Ampere, with a single H100 as the single-GPU deployment target. The weights, training data, and recipes are released under the OpenMDW-1.1 license. Full details are in NVIDIA's announcement and the model card.

Dedicated Endpoint

Docs

A Dedicated Endpoint is your own deployment of Nemotron 3.5 Lightning 30B A3B NVFP4, on an inference stack that Modal tunes and autoscales. Bring fine-tuned weights if you have them. Compute is billed by the second, and only while it runs.

$ modal endpoint create --model nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4

Related resources

Ship your first app in minutes.

Get Started

$30 / month free compute