Radeon 780M iGPU offers budget route for running MoE LLMs locally
Post claims AMD's Radeon 780M iGPU runs MoE models like Qwen 3.6 35B-A3B Q8 via partial expert offloading for better tokens/sec. Uses `--spec-type draft-mtp` speculative decoding with `--n-cpu-moe` to keep generation speed up while experts spill to CPU.
1 source
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- Gemma team to hold special event on August 20
- Situational Awareness invests $400M in chip startup Source Foundry
- Developer creates ChipTycoon to learn chip manufacturing via LLM simulation
- Reddit user creates AI-generated Big Bang Theory sitcom with ComfyUI
- Prompt injection is the most common way scammers attack people and agents