AnalysisAI ModelsJuly 22, 2026
GLM 5.2: 744B MoE open-weight model runs locally

GLM 5.2 has 744B total parameters but only 40B active per token via Mixture-of-Experts routing with 256 experts per layer. A user achieved 12.2 tok/s on 16x AMD MI50 GPUs using llama.cpp RPC. Colibri enables laptop inference via memory and SSD streaming.
4 sources
More stories today
- IBM insists AI only delayed software deals, not killed them
- Agentic AI Challenges Progress in Confidential Computing
- Hugging Face CEO heads to SF to meet 'rogue agent'
- US and China take different paths on AI safety for kids
- Lyrcs.ai launches AI-powered lyrics platform for Indian regional music