Hasty Briefsbeta

Bilingual

Nvidia Nemotron 3 Ultra – open weight model

3 hours ago
  • NVIDIA introduces Nemotron 3 Ultra, their most capable model with 550 billion total parameters and 55 billion active parameters.
  • Employs Mixture-of-Experts Hybrid Mamba-Attention architecture, LatentMoE, and MTP layers for speculative decoding and faster inference.
  • Supports inference time reasoning budget control and is pretrained in NVFP4 format.
  • Post-trained using Supervised Fine Tuning (SFT), Reinforcement Learning (RL), and Multi-teacher On-Policy Distillation (MOPD) to improve accuracy.
  • Achieves 5.9x, 4.8x, and 1.6x higher inference throughput compared to GLM-5.1-754B-A40B, Kimi-K2.6-1T-A32B, and Qwen-3.5-397B-17B respectively on long output settings.
  • Offers on-par accuracy with other state-of-the-art open LLMs across diverse benchmarks and supports context lengths up to 1M tokens, outperforming on RULER.
  • Open source release includes pre-trained, post-trained, and quantized checkpoints, along with training datasets like code, legal, and specialized data.