LaunchAI ModelsJune 26, 2026

NVIDIA Nemotron 3 Ultra NVFP4 checkpoint achieves 5.9x throughput boost

Quantized to NVFP4 using NVIDIA Model Optimizer, the 550B model achieves up to 5.9x higher inference throughput than GLM-5.1 754B FP4 while matching BF16 accuracy, and reduces model size from 1121 GB to 352.3 GB (3.2x). A single checkpoint runs on both Hopper (W4A16) and Blackwell (W4A4) hardware.

1 source

More stories today

Open the live feed