New papers tackle KV cache compression for LLM serving
Multiple arXiv papers propose KV cache compression methods: Minima-KV uses mixed-format paged attention, PuzzleKV applies page-wise low-rank decomposition, and KVBoost enables chunk-level reuse with deviation-guided recomputation. Others explore tradeoffs between tensor parallelism and KV compression, and question when attention is compressible.
How this story unfolded
2 days · 9 reports · from Aug 25
- Aug 25
- Aug 26
PuzzleKV: Page-Wise Low-Rank Decomposition for KV Cache Compressionarxiv.org
Minima-KV: Retention-Preserving KV Cache Compression with Mixed-Format Paged Attentionarxiv.org
VisCache: Visual KV Cache Pruning for Efficient Vision Large Language Model Inferencearxiv.org
More GPUs or a Smaller Cache? Tensor Parallelism versus KV Compression for Memory-Bound LLM Servingarxiv.org
- Aug 27
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- Redditor tests AI agents with $1 online task
- AI agents need their own identity before a gateway
- Claude Max users find default $200K spend limit
- TTFT-First Benchmark Ranks Lowest-Latency Voice and Realtime Agent APIs
- AI training demand causes Mac Mini shortages