AnalysisDevelopersSeptember 3, 2026

Qwen3.8-Flash-Next MTP support merged into ik_llama.cpp

Multi-token prediction support for Qwen3.8-Flash-Next landed on ik_llama.cpp main via PR #2369, doubling throughput from 45 to 90 tok/s on a 5090 with 128GB. It runs down to a 12GB 4070, and no fork or patch is needed.

1 source

More stories today

Open the live feed