Qwen 35B-A3B MoE runs ~4x faster than 27B dense locally
In local coding-maintenance tests on an R9700/llama.cpp setup, Qwen 35B-A3B MoE hit ~116 tok/s vs ~30 tok/s for the 27B dense model (~3.9x faster), while the coding-quality gap was much smaller than expected.
1 source
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- Anthropic's Ultracode coding mode gains industry attention
- Testing the Motion Context node for Stable Diffusion
- Krea2 Turbo BBOX fine-tune uploaded to HuggingFace
- Reddit user shares Minimax H3 character/object V2V swapping template
- ChatGPT accidentally makes photorealistic image mistaken for real photo