llama.cpp adds MTP support for GLM-4.5-Air

llama.cpp PR #26534 enables Multi-Token Prediction (MTP) for GLM-4.5-Air, a 106B MoE with 12B active parameters, offering speedups on memory-rich, compute-limited hardware like Strix Halo or DGX Spark.
1 source
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- Enterprise AI agents limited by messy documents
- Seinfeld AI video shows George in GTA 6 using Minimax H3
- Claude Code adds unrequested corrections to spec
- Ethan Mollick: AI impact research must address older-model limits
- Hobbyist trains 1.2B game music generator on single H100