AnalysisAI ModelsSeptember 12, 2026

Qwen3.8-27B-Q4 runs ~170k tokens on a 32 GB GPU without KV cache quantization

A LocalLLaMA user reports fitting roughly 170k tokens of Qwen3.8-27B-UD-Q4_K_XL on a 32 GB GPU with MTP and mmproj enabled, avoiding KV cache quantization entirely. They say even q8_0 KV cache quantization is perceptibly worse.

1 source

More stories today

Open the live feed