llama.cpp PR makes Q2_0 3.0–3.6x faster on x86 CPUs

llama.cpp PR #26348 adds an x86 VNNI implementation for the Q2_0 × Q8_0 dot product. The author's controlled CPU-only benchmarks show 3.0–3.6x higher throughput, with 8B decode rising from 2.39 to 8.20 tok/s.
1 source
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- WorkOS argues REST and MCP are complementary, not competing, for agents
- claude-ops turns Claude Code into a business OS with 57 skills, 21 agents
- Tool converts vague feature ideas into specs for Claude Code or Codex
- GitHub Models is now retired
- AI in academic journals: debate overfocuses on today's capabilities