LaunchAI ModelsJune 26, 2026
NVIDIA Nemotron 3 Ultra NVFP4 checkpoint achieves 5.9x throughput boost

Quantized to NVFP4 using NVIDIA Model Optimizer, the 550B model achieves up to 5.9x higher inference throughput than GLM-5.1 754B FP4 while matching BF16 accuracy, and reduces model size from 1121 GB to 352.3 GB (3.2x). A single checkpoint runs on both Hopper (W4A16) and Blackwell (W4A4) hardware.
1 source
More stories today
- Europe's Multilingual Reality Exposes AI Security Gaps
- Swiss AI releases Apertus-v1.5 8B and 70B language models
- Reddit: Distillation claims against China are 'ridiculous'
- AI researcher criticizes media for anthropomorphizing AI systems
- Alibaba reportedly tests standalone Qwen Office product