GLM-5.3 Flash launches, rivals Opus 4.8 at fraction of cost
Z.AI's GLM-5.3 Flash (formerly Ox Alpha) packs 320B parameters, 18B active, 1M context, and hybrid attention. On DeepSWE it nearly matches Luna's performance while doing twice as much work for the same budget. Unsloth released GGUF quants, and it runs on Databricks and Together AI.
How this story unfolded
3 days · 5 reports · 8 community posts · 13 of 14 shown
- Aug 27
- Aug 28
- Aug 29
- Aug 30
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- Google AI's EnvHarness turns static agent benchmarks into adaptive training worlds
- AI agents that pass authentication can still drift, expose data, or get memory-poisoned
- Local video watermark remover released on CivitAI
- Video walks through implementing Kimi K3 from scratch in PyTorch
- SeedVR2 TensorRT Studio offers free open-source video upscaling