AnalysisDevelopersAugust 27, 2026

llama.cpp PR caches hot MoE experts on GPU, boosting speed

A llama.cpp PR (#26563) caches frequently used MoE experts in VRAM while cold experts run on CPU, reporting 33→56 tok/s with 8GB VRAM. A user offloading hot experts to VRAM saw 50% boost (20→30 t/s) on Qwen 3.8 Flash Next.

How this story unfolded

4 weeks · 1 report · 2 community posts · from Aug 4

  1. Aug 4
  2. Aug 27
  3. Aug 29

Daily brief

Get tomorrow's AI brief in your inbox

More stories today

Open the live feed
llama.cpp PR caches hot MoE experts on GPU, boosting speed