Nvidia Nemotron 3 Ultra – open weight model
3 hours ago
- NVIDIA introduces Nemotron 3 Ultra, their most capable model with 550 billion total parameters and 55 billion active parameters.
- Employs Mixture-of-Experts Hybrid Mamba-Attention architecture, LatentMoE, and MTP layers for speculative decoding and faster inference.
- Supports inference time reasoning budget control and is pretrained in NVFP4 format.
- Post-trained using Supervised Fine Tuning (SFT), Reinforcement Learning (RL), and Multi-teacher On-Policy Distillation (MOPD) to improve accuracy.
- Achieves 5.9x, 4.8x, and 1.6x higher inference throughput compared to GLM-5.1-754B-A40B, Kimi-K2.6-1T-A32B, and Qwen-3.5-397B-17B respectively on long output settings.
- Offers on-par accuracy with other state-of-the-art open LLMs across diverse benchmarks and supports context lengths up to 1M tokens, outperforming on RULER.
- Open source release includes pre-trained, post-trained, and quantized checkpoints, along with training datasets like code, legal, and specialized data.