How-ToDevelopersAugust 9, 2026

Radeon 780M iGPU runs Qwen 3.6 35B-A3B via expert offloading

A Reddit user pitches the Radeon 780M iGPU as a budget option for MoE models like Qwen 3.6 35B-A3B Q8, using partial expert offloading (`--n-cpu-moe`) for a bigger tokens/s boost than full offload, plus MTP draft-mode speculation (`--spec-type draft-mtp --spec-draft-n-max 3`).

1 source

Daily brief

Get tomorrow's AI brief in your inbox

More stories today

Open the live feed