AnalysisAI ModelsJuly 22, 2026

GLM 5.2: 744B MoE open-weight model runs locally

GLM 5.2 has 744B total parameters but only 40B active per token via Mixture-of-Experts routing with 256 experts per layer. A user achieved 12.2 tok/s on 16x AMD MI50 GPUs using llama.cpp RPC. Colibri enables laptop inference via memory and SSD streaming.

4 sources

More stories today

Open the live feed