AnalysisAI ModelsJuly 22, 2026

GLM-5.2 Q4 runs at 12.2 tok/s on 16x AMD MI50

A Reddit user reports running GLM-5.2 UD-Q4_K_XL GGUF on 16x AMD MI50 32GB GPUs via llama.cpp RPC, achieving 12.2 tok/s output and 30.9 tok/s input at 10.7k context. Two parallel requests yielded 14.5 tok/s aggregate with 2x 16,384-token slots.

1 source

More stories today

Open the live feed