llama.cpp PR: up to 169% faster quantized-KV decode on Intel Battlemage

PR #26689 reports up to 169% faster decode at 118K context with q4_0/q8_0 quantized KV caches. The gain comes from switching a SYCL FlashAttention dispatch from the VEC kernel to TILE on Intel Battlemage.
1 source
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- Anthropic's Ultracode coding mode gains industry attention
- Testing the Motion Context node for Stable Diffusion
- Krea2 Turbo BBOX fine-tune uploaded to HuggingFace
- Reddit user shares Minimax H3 character/object V2V swapping template
- ChatGPT accidentally makes photorealistic image mistaken for real photo