How-ToDevelopersSeptember 29, 2026

Guide runs Qwen 3.8 27B Q4 at 100K context on 16GB AMD GPU

Read original source →reddit.com

A Reddit user reports ~30 t/s decode with Qwen 3.8 27B Q4 XS and q8/q5 KV cache on a 16GB RX 7800 XT, using llama.cpp built with Vulkan. The setup targets users who assumed the 27B model was infeasible on 16GB VRAM.

1 source

More stories today

Open the live feed