AnalysisAI ModelsJune 26, 2026

NVIDIA Nemotron 3 Ultra NVFP4 delivers up to 5.9x throughput, matches BF16 accuracy

The quantized Nemotron 3 Ultra NVFP4 achieves up to 5.9x higher inference throughput than GLM-5.1 754B FP4 on decode-heavy workloads while matching BF16 accuracy. Using mixed-precision NVFP4 quantization, the model shrinks from 1,121 GB BF16 to 352.3 GB and runs on both Hopper and Blackwell GPUs.

1 source

More stories today

Open the live feed