llama.cpp PR adds MTP support for Qwen3-Next

Pull request #25589 by yomaytk adds multi-token prediction (MTP) support for Qwen3-Next in llama.cpp, letting the model run 'at full speed.'
1 source
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- Post-training course materials invite educator feedback
- Kimi K3 available to try free on Together Chat
- Rhodium's Goujon urges holistic AI safety approach
- Cheap AI intelligence revives graph knowledge and ontologies
- US will exempt Chinese open-weight models from safety testing requirements