LaunchDevelopersJuly 29, 2026
llama.cpp loads MTP tensors by default even with MTP disabled

Recent llama.cpp builds load MTP/NextN tensors by default for draft-mtp architectures (GLM-5.2, hy_v3, qwen35moe, step35), even when speculative decoding is never enabled via --spec-type draft-mtp. Most community GGUFs bundle these layers, so users may see extra VRAM consumption unless the MTP weights are stripped from the file.
1 source
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- Semantica provides open-source enterprise intelligence layer for AI agents
- Best practices for creating professional-grade agent skills
- Creative Intelligence Suite provides agents for structured ideation
- LocalLLaMA community hyped over wave of mid-size model releases
- Peter Steinberger: 5.5 handles concurrent tasks without confusion