AnalysisAI ModelsJuly 8, 2026

4-bit GLM-5.2 (753B) scores 70.8% on Terminal-Bench 2.1

Quantized to Int4-Int8Mix with NVFP4 KV cache, GLM-5.2 (753B MoE) runs at TP=4 on 4× DGX Spark (GB10) with 100K context. It scores 63/89 (70.8%) on Terminal-Bench 2.1 vs 81.0% for the full model, using the same Terminus-2 agent scaffold.

1 source

Daily brief

Get tomorrow's AI brief in your inbox

More stories today

Open the live feed