AMD llama.cpp patch cuts MTP overhead, boosts Qwen 27B context to 149K
Reducing MTP buffer overhead in llama.cpp on AMD expanded Qwen 27B context from 64,256 to 149,248 tokens on a 16GB+12GB setup. The patch also lifted ROCm context on a single 16GB GPU from 19,456 to 76,032 tokens.
1 source
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- Claude Code 2.1.227 fixes subscription-tier, Bash and TUI bugs
- Curated resources for the open Agent2Agent protocol
- Suno to cap song downloads to curb AI slop
- Claude Code plugin translates 'Claudish' output into plain English
- Claude Code v2.1.227 fixes flag evaluation and Bash command failures