AnalysisAI ModelsSeptember 2, 2026

Q8 N-gram layer swapped into IQ4 Qwen 3.8 runs with no speed loss

A Reddit user reports replacing the 51B N-gram layer in Qwen 3.8 Next with a higher-precision Q8 version inside an IQ4 quant, with no speed degradation. A separate 5090 user reportedly swapped the N-gram portion of a Qwen 3.8 UD Q4 model to BF16.

1 source

More stories today

Open the live feed
Q8 N-gram layer swapped into IQ4 Qwen 3.8 runs with no speed loss