llama.cpp PR caches hot MoE experts on GPU, boosting speed

A llama.cpp PR (#26563) caches frequently used MoE experts in VRAM while cold experts run on CPU, reporting 33→56 tok/s with 8GB VRAM. A user offloading hot experts to VRAM saw 50% boost (20→30 t/s) on Qwen 3.8 Flash Next.
How this story unfolded
4 weeks · 1 report · 2 community posts · from Aug 4
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- Claude Code silently deletes local history older than 30 days
- Split sigmas speed up video reference in Stable Diffusion
- MiniMax tool turns anything into realistic human video
- MirroS' Code-as-World rewrites videos into executable MuJoCo programs
- PRAXIST research system boosts MLE-bench scores 44%