Nemotron 3 Super 120B A12B BF16 is a hybrid Mixture-of-Experts reasoning model developed by NVIDIA for large-scale agentic workflows and long-context reasoning. It features a 120B parameter architecture with ~12B active parameters, built on an 88-layer hybrid backbone combining Mamba-2 sequence layers, Transformer attention, and Latent MoE routing with 512 experts. The model incorporates Multi-Token Prediction (MTP) for faster generation and supports up to a 1M token context window, enabling efficient long-horizon reasoning, multi-agent coordination, and large-scale retrieval tasks.
Model specifications and capabilities are published by the model author and reproduced here from HF Model Card (nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16). Released under nvidia-open-model-license. Further reading: Paper · Blog
Choose your hardware and inference engine to get deployment commands and performance benchmarks tailored to your infrastructure
1,440 GB
Enter your email to get access to this content
Benchmarks measured by Vultr on B200 · vLLM. Throughput and latency vary with concurrency, input length and engine configuration, so treat these as a comparison baseline, not a service guarantee.