AnalysisAI ModelsJuly 3, 2026
GLM-5.2 NVFP4 hits ~24 tok/s on four DGX Sparks at 128K context

An r/LocalLLaMA benchmark reports GLM-5.2 NVFP4 now runs at ~24 tok/s at 128K context across four DGX Spark units. The post says the MTP slowdown that previously capped throughput at ~15 tok/s is solved.
1 source
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- Semantica provides open-source enterprise intelligence layer for AI agents
- Best practices for creating professional-grade agent skills
- Creative Intelligence Suite provides agents for structured ideation
- LocalLLaMA community hyped over wave of mid-size model releases
- Peter Steinberger: 5.5 handles concurrent tasks without confusion