AnalysisAI ModelsJuly 22, 2026
GLM-5.2 Q4 runs at 12.2 tok/s on 16x AMD MI50
A Reddit user reports running GLM-5.2 UD-Q4_K_XL GGUF on 16x AMD MI50 32GB GPUs via llama.cpp RPC, achieving 12.2 tok/s output and 30.9 tok/s input at 10.7k context. Two parallel requests yielded 14.5 tok/s aggregate with 2x 16,384-token slots.
1 source
More stories today
- Worker building AI data center says budgets are 'unreal'
- CARPL.ai Raises $10M to Expand Radiology AI Marketplace
- Dynatrace announces deterministic SRE agents for AI ops
- Cloudflare open-sources privacy protocol debugger for AI agents
- SSI and Nvidia announce strategic partnership and investment