LaunchDevelopersJuly 29, 2026

llama.cpp now loads MTP tensors by default for draft-mtp models

Recent llama.cpp builds load MTP/NextN tensors by default for any draft-mtp architecture, even when speculative decoding is disabled. Previously these tensors were skipped unless users enabled it explicitly; GGUFs bundling MTP weights (GLM-5.2, hy_v3, qwen35moe, step35) are affected.

1 source

Daily brief

Get tomorrow's AI brief in your inbox

More stories today

Open the live feed