LaunchAI ModelsJuly 9, 2026
NVIDIA releases Nemotron-Labs-3-Puzzle-75B-A9B compressed MoE

The compressed variant of Nemotron-3-Super achieves 2.03x server throughput at matched user throughput. It uses active parameters, KV cache, and Mamba state to serve more users per node.