LaunchAI ModelsJuly 9, 2026

NVIDIA releases compressed MoE model Nemotron-Labs-3-Puzzle-75B-A9B

The compressed hybrid MoE variant achieves 2.03x server throughput at matched user throughput compared to the original Nemotron-3-Super. It reduces active parameters, KV cache, and Mamba state to improve serving efficiency.

1 source

More stories today

Open the live feed