AnalysisAI ModelsJuly 8, 2026
4-bit GLM-5.2 (753B) scores 70.8% on Terminal-Bench 2.1
Quantized to Int4-Int8Mix with NVFP4 KV cache, GLM-5.2 (753B MoE) runs at TP=4 on 4× DGX Spark (GB10) with 100K context. It scores 63/89 (70.8%) on Terminal-Bench 2.1 vs 81.0% for the full model, using the same Terminus-2 agent scaffold.
1 source
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- Semantica provides open-source enterprise intelligence layer for AI agents
- Best practices for creating professional-grade agent skills
- Creative Intelligence Suite provides agents for structured ideation
- LocalLLaMA community hyped over wave of mid-size model releases
- Peter Steinberger: 5.5 handles concurrent tasks without confusion