AnalysisAI ModelsAugust 8, 2026

Qwen 35B-A3B MoE runs ~4× faster than 27B dense in local coding tests

In local coding tests on an R9700/llama.cpp rig, the MoE hit ~116 tok/s vs ~30 tok/s (~3.9× faster), yet coding-quality trailed the dense 27B far less than expected, with both handling ordinary bug fixes similarly.

1 source

Daily brief

Get tomorrow's AI brief in your inbox

More stories today

Open the live feed