AnalysisAI ModelsJune 26, 2026
NVIDIA Nemotron 3 Ultra NVFP4 delivers up to 5.9x throughput, matches BF16 accuracy

The quantized Nemotron 3 Ultra NVFP4 achieves up to 5.9x higher inference throughput than GLM-5.1 754B FP4 on decode-heavy workloads while matching BF16 accuracy. Using mixed-precision NVFP4 quantization, the model shrinks from 1,121 GB BF16 to 352.3 GB and runs on both Hopper and Blackwell GPUs.
1 source
More stories today
- Robotics startup Genesis AI in talks to raise $500M
- Cal AI founder discusses ROI of MrBeast partnership
- Sonilo launches Sound Effects 1.0 audio model
- Nvidia signs $1.5B chip packaging deal with Amkor
- AI for the Aging Population