LaunchDevelopersJuly 29, 2026

llama.cpp loads MTP tensors by default even with MTP disabled

Recent llama.cpp builds load MTP/NextN tensors by default for draft-mtp architectures (GLM-5.2, hy_v3, qwen35moe, step35), even when speculative decoding is never enabled via --spec-type draft-mtp. Most community GGUFs bundle these layers, so users may see extra VRAM consumption unless the MTP weights are stripped from the file.

1 source

Daily brief

Get tomorrow's AI brief in your inbox

More stories today

Open the live feed
llama.cpp loads MTP tensors by default even with MTP disabled — AIBriefs