Z.ai releases GLM-5.3-Flash, a 320B-A18B multimodal model

GLM-5.3-Flash is a 320B-A18B natively multimodal model with a 1M-token context window, released under the MIT License. It scores 57 on the Artificial Analysis Intelligence Index at $0.09 cost per task, with pricing at $0.15 per million input tokens and $0.50 output.
How this story unfolded
4 days · 13 reports · 30 community posts · 43 of 49 shown
- Aug 25
- Aug 26
China’s Z.AI Made Ox Alpha Stealth Model That Rivals DeepSeekbloomberg.com
Surprise: Z.ai is the AI lab behind the mysterious Ox Alpha modeltechcrunch.com
zai-org/GLM-5.3-Flashhuggingface.co
GLM 5.3 Flash now available on AI Gatewayvercel.com
Z.ai’s GLM-5.3 Flash is cheap, good, and served on Chinese chipsthenewstack.io
unsloth/GLM-5.3-Flash-GGUFhuggingface.co
Z.ai Releases GLM-5.3-Flash: A 320B-A18B Natively Multimodal MoE With a 1M-Token Contextmarktechpost.com
- Aug 27
- Aug 28
Z.ai by email
Get an email when Z.ai has news
No news that day, no email.
More stories today
- Musicians-turned-detectives hunt AI-generated music grifters
- Krea 2 (Roma) macro workflow in Nomad Studio
- Reddit user observes speculative decoding at low t/s
- llmog: local LLM tool for auto-annotating datasets
- David Ha: model resiliency key as coding tools lose frontier access