AnalysisDevelopersOctober 1, 2026

llama.cpp adds MTP and halves indexer memory for Qwen Flash Next

Read original source →github.com

Two llama.cpp PRs land for Qwen Flash Next: MTP support (merged after 17h of development) and an indexer-score memory cut that halves VRAM use. GGUF quants are published at ggml-org/Qwen3.8-Flash-Next-GGUF.

2 sources

More stories today

Open the live feed