llama.cpp PR makes Q2_0 3.0–3.6x faster on x86 CPUs

The pull request adds an x86 VNNI kernel for Q2_0×Q8_0 dot products, with author benchmarks showing roughly 3–3.6x higher throughput. In one test, an 8B model's CPU decode speed rose from 2.39 to 8.20 tok/s.
1 source
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- Anthropic's Ultracode coding mode gains industry attention
- Testing the Motion Context node for Stable Diffusion
- Krea2 Turbo BBOX fine-tune uploaded to HuggingFace
- Reddit user shares Minimax H3 character/object V2V swapping template
- ChatGPT accidentally makes photorealistic image mistaken for real photo