How-ToDevelopersSeptember 6, 2026

Optimized Qwen3.8 27B stack released for Strix Halo

A reproducible ROCm stack for Qwen3.8-27B on Ryzen AI Max / Max+ (gfx1151) with calibrated IQ4_XS weights, an IQ4_XS DFlash2 drafter, and retained-PM4 dispatch. Benchmarked on a Radeon 8060S with a 31,497-token prompt and 256-token output, the custom ROCm build claims fastest prefill and decode.

1 source

More stories today

Open the live feed
Optimized Qwen3.8 27B stack released for Strix Halo