AnalysisAI ModelsSeptember 3, 2026

Runtime router tweak cuts reasoning tokens 8.5% in MoE models

A new paper shows that expanding the expert selection budget at runtime, without retraining, reduces reasoning token usage by 8.5% in sparse MoE models like Qwen 35B A4B+.

1 source

More stories today

Open the live feed