llama.cpp PR speeds quantized-KV decode up to 169% on Intel Battlemage

PR #26689 switches llama.cpp's SYCL FlashAttention decode path from the VEC to TILE kernel for quantized KV caches (q4_0/q8_0), delivering up to 169% faster decode at 118K context on Intel Battlemage.
1 source
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- Vercel launches eve, its 'Next.js for agents' framework
- Black Hat 2026 wrap-up stresses open frameworks for agentic security
- Together AI adds educational documentation for LLM development concepts
- Auto mode launches after many months of internal use
- JPMorgan sees tech bond sales topping $500B as AI debt binge expands