ExFold paper and llama.cpp PR speed up MoE by caching hot experts

ExFold (arXiv:2608.24938) accelerates MoE prefill-decode without training. A llama.cpp PR (#26563) caches frequently used experts in VRAM, reporting 33→56 tok/s on 8GB VRAM; a user saw 20→30 t/s on Qwen 3.8 Flash Next.
How this story unfolded
4 weeks · 1 report · 2 community posts · from Aug 4
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- AI band gets YouTube Official Artist Channel status
- Sony Music, Warner sue Anthropic over alleged IP theft
- AWS engineer shows robot answering untrained questions
- Bernie Sanders vows legislation to stop Flock AI surveillance
- Claude-built market-timing game: couch beats players 62%