llama.cpp PR makes Q2_0 up to 3.6x faster on x86 CPUs

llama.cpp PR #26348 adds an x86 VNNI implementation for the Q2_0 × Q8_0 dot product, delivering 3.0–3.6x higher throughput in CPU-only benchmarks; an 8B model's decode speed jumps from 2.39 to 8.20 tok/s.
1 source
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- a16z podcast: AI models now exploit vulnerabilities, not just find them
- Claude Opus 5 praised by student, then fails simple chart task
- Airbnb tests AI-powered search with user-controlled toggle
- Bernie Sanders worries AI is a 'Don't Look Up' situation
- Reddit user spent $3,000 building a dream game with Claude Code