How-ToAI ModelsJuly 8, 2026
GLM 5.2 runs on 4xGB10 cluster at ~25 t/s with 330k context

Post details running GLM 5.2 across 4 GB10s on a 100G switch with TP4+DCP2, hitting ~25 tokens/s decode and ~650 tokens/s prefill at ~330k context. Using 4 draft tokens (vs Z.ai's rec of 5); dropping DCP raises prefill to ~750 t/s, and DCP4 extends to 660k context at ~400 t/s prefill.
1 source
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- Semantica provides open-source enterprise intelligence layer for AI agents
- Best practices for creating professional-grade agent skills
- Creative Intelligence Suite provides agents for structured ideation
- LocalLLaMA community hyped over wave of mid-size model releases
- Peter Steinberger: 5.5 handles concurrent tasks without confusion