llama.cpp PR adds GPU caching for frequently used MoE experts
A new pull request (#26563) implements a heatmap to track and cache 'hot' MoE experts in VRAM, leaving cold experts on the CPU. Users report performance gains from 33 to 56 tokens per second on 8GB VRAM configurations.
1 source
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- Koharu translates manga pages locally using OCR, inpainting, and LLMs
- Orchestration tool runs AI coding agents in parallel, compares answers
- Real-time Gaussian splats generated on mobile devices
- Newtake AI shows off completely AI-generated rap music video
- ChatGPT Work used to install OpenClaw and Ollama, run local model