Qwen 35B-A3B MoE ~4× faster than 27B dense, small coding-quality gap
In local coding-maintenance tests, the MoE model hit ~116 tok/s vs ~30 tok/s for the dense model — a 3.9× speedup. The quality gap on routine bug-fixes was much smaller than expected, with both handling ordinary tasks well.
1 source
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- Koharu translates manga pages locally using OCR, inpainting, and LLMs
- Orchestration tool runs AI coding agents in parallel, compares answers
- Real-time Gaussian splats generated on mobile devices
- Newtake AI shows off completely AI-generated rap music video
- ChatGPT Work used to install OpenClaw and Ollama, run local model