AnalysisAI ModelsJuly 22, 2026
GLM 5.2: How 744B-parameter model runs with 40B active

GLM 5.2 has 744B total parameters but only 40B active per token via Mixture of Experts routing, enabling inference on consumer hardware. The architecture uses 256 experts per layer and can be run locally using Colibri's three-tier memory and SSD streaming.
3 sources
More stories today
- Vercel co-signs Open Weights and American AI Leadership letter
- How Autonomous AI Is Transforming Chip and System Design
- NVIDIA expands Agent Toolkit with PhysicsNeMo and CUDA-X libraries
- NVIDIA advances semiconductor materials engineering and manufacturing
- NVIDIA Nemotron 3 Ultra achieves benchmark-leading performance with LangChain Deep Agents