AnalysisAI ModelsSeptember 12, 2026

Qwen3.8-27B-Q4 runs at ~170k context on a 32 GB GPU without KV cache quantization

A LocalLLaMA user reports fitting roughly 170k tokens of Qwen3.8-27B-UD-Q4_K_XL on a 32 GB GPU with MTP and mmproj enabled, avoiding KV cache quantization entirely. They say even q8_0 KV cache quantization is noticeably worse, and Qwen3.8 burns through the context at xhigh effort.

1 source

More stories today

Open the live feed