Nemotron 3 Super is NVIDIA's 120B-parameter latent Mixture-of-Experts model with 12B active per token, designed for multi-agent reasoning, coding, and long-context work. It takes text input, supports up to 1M tokens of context and configurable thinking, and is released under the NVIDIA Nemotron Open Model License.
Nemotron 3 Super pairs a hybrid Mamba-2 and Transformer backbone with a latent Mixture-of-Experts design: tokens are projected into a smaller latent dimension before expert routing, letting the model consult four times as many experts for the same compute. The network has 88 layers, 512 routed experts plus one shared expert with 22 active per token, and 120 billion total parameters of which 12 billion are active. It was pre-trained on 25 trillion tokens with most linear layers running in NVFP4, and includes shared-weight multi-token prediction heads for native speculative decoding. Context extends up to 1 million tokens, thinking can be switched on or off via the chat template, and it supports English, French, German, Italian, Japanese, Spanish, and Chinese.
This is NVIDIA's NVFP4 release of Nemotron 3 Super, the first model in the Nemotron 3 family trained at that precision rather than quantized afterwards. NVIDIA lists a single B200 as the minimum GPU requirement. The weights are released under the NVIDIA Nemotron Open Model License. Full details are in NVIDIA's announcement and the model card.
A Dedicated Endpoint is your own deployment of NVIDIA Nemotron 3 Super 120B A12B NVFP4, on an inference stack that Modal tunes and autoscales. Bring fine-tuned weights if you have them. Compute is billed by the second, and only while it runs.
modal endpoint create --model nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4