llama.cpp PR caches hot MoE experts on GPU, boosting speed
A llama.cpp PR (#26563) adds a heatmap to track frequently used MoE experts, caching them in VRAM while cold experts run on CPU. The author reports 33→56 tok/s on 8GB VRAM.
1 source
Developers by email
Get an email when there's news on Developers
No news that day, no email.
More stories today
- Open-source RL training with trl and OpenEnv shared
- NVIDIA invests $3.5B in MediaTek, deepens AI partnership
- Neta team explains why their open-source model generated Anne Hathaway-like images
- OpenAI age-verification error deletes adult's account
- South Korea gives citizens free unlimited domestic AI access