LaunchAI ModelsJune 26, 2026
NVIDIA releases Nemotron 3 Ultra NVFP4 checkpoint with up to 5.9x throughput gain

The NVFP4 quantized checkpoint achieves up to 5.9x higher inference throughput than GLM-5.1 754B FP4 on decode-heavy workloads while matching BF16 accuracy. Model size is reduced from 1121 GB to 352.3 GB (3.2x reduction), and the checkpoint can run on both Hopper and Blackwell hardware by automatically converting weight format.
1 source
More stories today
- Hugging Face Uses Chinese Model in Self-Defense
- Claude replays all players in daily racing game Swervle
- How AI guardrails are impeding the work of offensive cybersecurity researchers
- Fugu-Ultra v1.1 beats Fable 5 by 7.9 points in coding and reasoning
- audio.cpp 0.4 adds Higgs Audio v3 TTS 4B and Fish Audio S2 Pro