AnalysisAI ModelsAugust 27, 2026

ExFold paper and llama.cpp PR speed up MoE by caching hot experts

ExFold (arXiv:2608.24938) accelerates MoE prefill-decode without training. A llama.cpp PR (#26563) caches frequently used experts in VRAM, reporting 33→56 tok/s on 8GB VRAM; a user saw 20→30 t/s on Qwen 3.8 Flash Next.

How this story unfolded

4 weeks · 1 report · 2 community posts · from Aug 4

  1. Aug 4
  2. Aug 27
  3. Aug 29

Daily brief

Get tomorrow's AI brief in your inbox

More stories today

Open the live feed