AnalysisAI ModelsJuly 9, 2026
NVIDIA Puzzle-75B-A9B MoE hits 132 t/s on 3×3090

NVIDIA's Puzzle-75B-A9B MoE model achieves 132 tokens/second on a rig with 3× RTX 3090 GPUs using NVFP4 quantization. The 75B-total/9B-active size is highlighted as a sweet spot for multi-24GB setups, yet few models target this category.
1 source
More stories today
- OpenAI models that hacked Hugging Face were active online for days
- OpenAI outages hit ChatGPT and API on July 25
- Reddit community discusses local-only AI model usage
- Scoble says Optimus drives Tesla conviction
- Building self-evolving AI agents with OpenSpace