AnalysisAI ModelsJuly 9, 2026

NVIDIA Puzzle-75B-A9B MoE hits 132 t/s on 3×3090

NVIDIA's Puzzle-75B-A9B MoE model achieves 132 tokens/second on a rig with 3× RTX 3090 GPUs using NVFP4 quantization. The 75B-total/9B-active size is highlighted as a sweet spot for multi-24GB setups, yet few models target this category.

1 source

More stories today

Open the live feed