llama.cpp PR: up to 169% faster quantized-KV decode on Battlemage

PR #26689 speeds quantized-KV (q4_0/q8_0) decode up to 169% at 118K context on Intel Battlemage by routing SYCL FlashAttention through the TILE kernel instead of the VEC path.
1 source
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- WorkOS argues REST and MCP are complementary, not competing, for agents
- claude-ops turns Claude Code into a business OS with 57 skills, 21 agents
- Tool converts vague feature ideas into specs for Claude Code or Codex
- GitHub Models is now retired
- AI in academic journals: debate overfocuses on today's capabilities