Nemotron 3 Ultra is NVIDIA's frontier-scale open model: a 550B-parameter latent Mixture-of-Experts with 55B active per token, built for frontier reasoning and orchestration in long-running agents. It takes text input and offers up to 1M-token context, configurable thinking, ten languages, and OpenMDW-1.1 weights.
Nemotron 3 Ultra scales the Nemotron 3 hybrid design to 550 billion total parameters with 55 billion active per token. Its 108 layers interleave 48 Mamba-2 layers, 48 latent Mixture-of-Experts layers with 512 routed experts plus one shared expert and 22 active per token, and 12 attention layers, so Mamba handles long sequences efficiently while attention preserves precise recall. Pre-training covered roughly 20 trillion tokens using an NVFP4 recipe, followed by supervised fine-tuning, asynchronous GRPO reinforcement learning, and Multi-Domain On-Policy Distillation from specialized teacher models. Multi-token prediction heads enable speculative decoding, context runs up to 1 million tokens, thinking is configurable via the chat template, and it supports ten languages including English, Chinese, Japanese, Korean, Hindi, and Brazilian Portuguese.
This is NVIDIA's own NVFP4 release of Nemotron 3 Ultra, trained with an NVFP4 pre-training recipe. It runs on Hopper, Blackwell, and Ampere GPUs via specialized NVFP4 kernels, and NVIDIA lists 4x B200, 4x GB200, 4x B300, 4x GB300, or 8x H100 as the minimum hardware. With Ultra, Nemotron releases move to the Linux Foundation's OpenMDW-1.1 license. Full details are in NVIDIA's announcement and the model card.
A Dedicated Endpoint is your own deployment of NVIDIA Nemotron 3 Ultra 550B A55B NVFP4, on an inference stack that Modal tunes and autoscales. Bring fine-tuned weights if you have them. Compute is billed by the second, and only while it runs.
modal endpoint create --model nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-NVFP4