AnalysisAI ModelsJuly 22, 2026

GLM 5.2: How 744B-parameter model runs with 40B active

GLM 5.2 has 744B total parameters but only 40B active per token via Mixture of Experts routing, enabling inference on consumer hardware. The architecture uses 256 experts per layer and can be run locally using Colibri's three-tier memory and SSD streaming.

3 sources

More stories today

Open the live feed