AnalysisDevelopersJuly 29, 2026

llama.cpp now loads MTP tensors by default even with MTP disabled

Recent llama.cpp builds load MTP/NextN tensors baked into GGUFs by default for draft-mtp architectures (GLM-5.2, hy_v3, qwen35moe, step35), even when speculative decoding is never enabled. Previously these tensors were skipped unless --spec-type draft-mtp was passed.

1 source

Daily brief

Get tomorrow's AI brief in your inbox

More stories today

Open the live feed