llama.cpp now loads MTP tensors by default for draft-mtp models

Recent llama.cpp builds load MTP/NextN tensors by default for any draft-mtp architecture, even when speculative decoding is disabled. Previously these tensors were skipped unless users enabled it explicitly; GGUFs bundling MTP weights (GLM-5.2, hy_v3, qwen35moe, step35) are affected.
1 source
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- South Korea's KOMCA reverses ban on AI-assisted music copyright
- Sakana AI, Daiwa Securities joint AI project enters full-scale production
- Multi-datacenter architecture gives Together AI 99.9% uptime
- LLM 0.32 adds reasoning traces, OpenAI Responses, server-side tools
- ACM Queue examines eight myths on software engineering and GenAI