AnalysisAI ModelsJuly 11, 2026
Qwen3-30B-A3B runs at 50 tok/s on RTX 5060 Ti with custom code

Custom CUDA/C++ code achieves 50-54 tok/s on Qwen3-30B-A3B at float8 on an RTX 5060 Ti with 16 GB VRAM, a 50% improvement over llama.cpp's 33-34 tok/s. The optimizations combine SOTA techniques for efficient MoE inference.