AnalysisAI ModelsAugust 27, 2026

GLM-5.3-Flash hits ~206 tok/s on DGX Station GB300

A Reddit user reports running GLM-5.3-Flash on an NVIDIA DGX Station GB300 at ~206 tokens/s (single stream) with 1M context. The user was experimenting with GLM-5.2 in anticipation of GLM-5.3's release.

1 source

More stories today

Open the live feed