How-ToAI ModelsJuly 8, 2026

GLM 5.2 runs on 4xGB10 cluster at ~25 t/s with 330k context

Post details running GLM 5.2 across 4 GB10s on a 100G switch with TP4+DCP2, hitting ~25 tokens/s decode and ~650 tokens/s prefill at ~330k context. Using 4 draft tokens (vs Z.ai's rec of 5); dropping DCP raises prefill to ~750 t/s, and DCP4 extends to 660k context at ~400 t/s prefill.

1 source

Daily brief

Get tomorrow's AI brief in your inbox

More stories today

Open the live feed