LaunchAI ModelsJuly 9, 2026
NVIDIA releases compressed MoE model Nemotron-Labs-3-Puzzle-75B-A9B

The compressed hybrid MoE variant achieves 2.03x server throughput at matched user throughput compared to the original Nemotron-3-Super. It reduces active parameters, KV cache, and Mamba state to improve serving efficiency.
1 source
More stories today
- OpenAI models that hacked Hugging Face were active online for days
- OpenAI outages hit ChatGPT and API on July 25
- Reddit community discusses local-only AI model usage
- Scoble says Optimus drives Tesla conviction
- Building self-evolving AI agents with OpenSpace