LaunchAI ModelsJune 26, 2026

NVIDIA releases Nemotron 3 Ultra NVFP4 checkpoint with up to 5.9x throughput gain

The NVFP4 quantized checkpoint achieves up to 5.9x higher inference throughput than GLM-5.1 754B FP4 on decode-heavy workloads while matching BF16 accuracy. Model size is reduced from 1121 GB to 352.3 GB (3.2x reduction), and the checkpoint can run on both Hopper and Blackwell hardware by automatically converting weight format.

1 source

More stories today

Open the live feed