llama.cpp PR caches hot MoE experts on GPU — 33 to 56 tok/s on 8GB VRAM
llama.cpp PR #26563 adds an expert-use heatmap that caches frequently used MoE experts in VRAM while cold experts stay on CPU; the author reports 33 → 56 tok/s on an 8GB VRAM system.
1 source
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- Post-training course materials invite educator feedback
- Kimi K3 available to try free on Together Chat
- Rhodium's Goujon urges holistic AI safety approach
- Cheap AI intelligence revives graph knowledge and ontologies
- US will exempt Chinese open-weight models from safety testing requirements