AnalysisDevelopersOctober 1, 2026

llama.cpp adds MTP and halves indexer memory for Qwen Flash Next

Read original source →github.com

Two llama.cpp PRs land for Qwen Flash Next: MTP support (merged after 17h of development) and a change halving indexer score memory to cut VRAM use. GGUF quants are published at ggml-org/Qwen3.8-Flash-Next-GGUF.

2 sources

More stories today

Open the live feed