Radeon 780M iGPU runs Qwen 3.6 35B-A3B via MoE expert offloading
Partial expert offloading for MoE models on the 780M gives a solid token-generation boost but slower prompt processing, per the r/LocalLLaMA post. A Qwen 3.6 35B-A3B Q8 setup uses --spec-type draft-mtp and --n-cpu-moe in llama.cpp.
1 source
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- Koharu translates manga pages locally using OCR, inpainting, and LLMs
- Orchestration tool runs AI coding agents in parallel, compares answers
- Real-time Gaussian splats generated on mobile devices
- Newtake AI shows off completely AI-generated rap music video
- ChatGPT Work used to install OpenClaw and Ollama, run local model