llama.cpp PR caches hot MoE experts on GPU — 33→56 tok/s with 8GB VRAM
PR #26563 adds an expert-usage heatmap to llama.cpp, caching frequently used MoE experts in VRAM while cold experts run on CPU. The author reports 33 → 56 tok/s on an 8GB VRAM setup.
1 source
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- AI detectors face criticism over reliability and impact on trust
- Student uses $3 chip to run Claude Code for automated betting
- MiniMax H3 CLIP swap cuts VRAM from 15.7 GB to 4.5 GB
- Artist's AI-generated 'Found [You?]' footage project blends video and music
- Anthropic's Haiku 4.5 nears 12 months without an update