AnalysisAI ModelsJuly 11, 2026
Running Qwen3 30B A3B at 50 tok/s on RTX 5060 Ti

Custom CUDA/C++ code achieves 50-54 tok/s for Qwen3-30B-A3B at float 8 on 16GB RTX 5060 Ti, a ~50% improvement over llama.cpp's 33-34 tok/s.
1 source
More stories today
- Circular financing ain't what it used to be
- Moonshot AI open-sources MoonEP, a communication library for distributed MoE.
- Patrick Lo warns against rushed AI adoption in healthcare
- Postdoc position open for HCI expert on governance project
- Apple Will 'Watch Everything Burn' When the AI Bubble Bursts