AnalysisDevelopersJuly 22, 2026

VRAM disk cache of MoE hits 340 pp/s for Kimi 2.7 on single DGX Spark

A llama.cpp trick using VRAM as disk cache reaches 340 pp/s and 9.6 tg/s running a 204GB Kimi-K2.7-Code GGUF on a single DGX Spark. The strategy keeps MoE experts on the CUDA compute path by caching them in VRAM over disk.

1 source

Daily brief

Get tomorrow's AI brief in your inbox

More stories today

Open the live feed