AnalysisDevelopersJuly 22, 2026
VRAM disk cache of MoE hits 340 pp/s for Kimi 2.7 on single DGX Spark

A llama.cpp trick using VRAM as disk cache reaches 340 pp/s and 9.6 tg/s running a 204GB Kimi-K2.7-Code GGUF on a single DGX Spark. The strategy keeps MoE experts on the CUDA compute path by caching them in VRAM over disk.
1 source
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- Thoughtworks' Kief Morris: humans must stay 'on the loop' in AI delivery
- GEMA wins major copyright ruling against Suno, orders damages paid
- LangChain builds ReviewBench benchmark for code review agents
- DeepSeek Flash 0731's reasoning trace amuses with 'OH MY GOD' outburst
- Former OpenAI VP Jerry Tworek discusses AI lab automation