llama.cpp now loads MTP tensors by default even with MTP disabled

Recent llama.cpp builds load MTP/NextN tensors baked into GGUFs by default for draft-mtp architectures (GLM-5.2, hy_v3, qwen35moe, step35), even when speculative decoding is never enabled. Previously these tensors were skipped unless --spec-type draft-mtp was passed.
1 source
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- Onton releases Ontology 1, a neurosymbolic search model
- Agentic SOC Platform uses AI agents for security triage
- WiFi-3D-Fusion performs real-time 3D human pose estimation
- Eight AI agents automate Obsidian vault
- Comfyanon says H3 will still release