AnalysisAI ModelsJuly 9, 2026

NVIDIA Puzzle-75B-A9B runs at 132 t/s on 3x3090 with NVFP4

User reports 132 tokens/second on 3x3090 GPUs using NVFP4 quantization and vLLM 0.22.1. The 75B parameter MoE model has 9B active parameters, making it efficient for multi-24GB setups.

1 source