AnalysisAI ModelsJuly 22, 2026
GLM 5.2 open-weight model runs 744B parameters on consumer hardware

GLM 5.2 has 744B total parameters but only 40B active per token via Mixture-of-Experts routing. Techniques like Colibri's three-tier memory streaming and llama.cpp RPC enable local inference on consumer laptops and multi-GPU setups.
4 sources
More stories today
- Fugu-Ultra v1.1 beats Fable 5 by 7.9 points in coding and reasoning
- Meta's AI optimism ad uses David Bowie song about apocalypse
- Amazon cracks down on use of AI images by sellers after New York law
- Claude zoom tool cookbook updated for high-res images
- User lets Claude direct a short movie