Z.ai releases open-weight GLM-5.3, beating rivals on agentic benchmarks

Z.ai released GLM-5.3 as open weights on Hugging Face, with GLM-5.3-Flash also available. It scores 84.5% on CyberGym and 60 on the Artificial Analysis Intelligence Index, beating GPT-5.6 Sol and Claude Fable 5 on agentic benchmarks.
How this story unfolded
3 weeks · 15 reports · 37 community posts · 52 of 55 shown
- Aug 14
Z.ai Ships GLM-5.3 Without Retraining the Base Model: Better at Complex Coding and Long-Horizon Tasksmarktechpost.com
GLM-5.3 didn’t change the base model — where did its coding gains come from?thenewstack.io
China's Z.AI Ships GLM-5.3, Calling It the Top Open-Weight Coding Modeldecrypt.co
GLM-5.3: How Chinese labs keep stride with the frontierinterconnects.ai
GLM-5.3 is here with advanced cyber capabilities — and reportedly already found a 'serious vulnerability' in Cursorventurebeat.com
- Aug 15
- Aug 16
- Aug 17
- Aug 18
The Powerful Chinese Model Experts Warned About—and Waited for—Is Herewired.com
OpenAI’s Greg Brockman: Z.ai’s GLM-5.3 likely to “significantly accelerate the threat landscape”thenewstack.io
GLM 5.3 now available on AI Gatewayvercel.com
GLM-5.3 achieves 60 on the Artificial Analysis Intelligence Index, on par with Kimi K3 and up 7...
- Aug 19
GLM-5.3 hits the API at $1.4/$4.4 per million tokensventurebeat.com
An industrial-scale distillation of models, or subtle benchmaxxing: What developers really think of GLM-5.3thenewstack.io
GLM-5.3 achieves 60 on the Artificial Analysis Intelligence Index, on par with Kimi K3 and up 7 points from GLM-5.2. Once the weights are released it will be tied as the leading open weights model
- Aug 23
- Aug 24
- Aug 28
- Aug 29
- Aug 30
- Aug 31
- Sep 1
- Sep 2
More stories today
Bill Gurley: Hold AI companies accountable for product behavior
Bill Gurley·2 hours agoMinimax H3 users seek faster generation methods
Reddit users discuss optimizing Minimax H3 video generation speed, comparing turbo LoRAs, SLA attention, and step counts. One user reports 1920x1088 10-second clips in 5 minutes using a 4-step turbo LoRA with SLA attention on an RTX 5090.
r/StableDiffusion·3 hours agoPressure sensors improve robotic gripping accuracy
Robotic gripping fails due to lack of real-time contact feedback, not mechanical strength. Pressure sensors measure distributed stress at contact, enabling early detection of micro slips and load redistribution within milliseconds.
The Robot Report·3 hours ago

Z.ai by email
Get an email when Z.ai has news
No news that day, no email.
Hermes Agent adds per-model provider pinning for OpenRouter
Teknium·4 hours ago
Reddit compares Fable 5.1 and Astra joke-writing
A Reddit user asked Fable 5.1 and Astra to write their funniest original joke, sparking a 78-comment debate. The post includes jokes from both models, with users voting on which is funnier.
r/ClaudeAI·4 hours agoUsers report GPT-5.6 Sol quality shifts after updates
Reddit users report GPT-5.6 Sol in ChatGPT feels 'lobotomized' or 'nerfed' since the 08/06/2026 update and Astra's release, while others claim it was secretly upgraded. Complaints cite reduced reasoning depth and laziness on coding and research tasks.
r/ChatGPT·4 hours ago
Block KV cache streaming bounds VRAM at long context
A pull request to llama-cpp-turboquant introduces block KV cache streaming via a shared CUDA phase arena, bounding VRAM usage at long context. The author ported and extended Raymond's work to multiple models beyond Qwen, benchmarking to confirm value.
r/LocalLLaMA·4 hours ago
Boris Cherny discusses Claude writing its own code
Boris Cherny, Head of Claude Code at Anthropic, discusses how Claude writes its own code. He previously built one of the fastest-growing developer tools and was a senior engineer at Meta.
YouTube·4 hours ago