AnalysisAI ModelsSeptember 3, 2026

Runtime router tweak cuts MoE reasoning tokens 8.5%

A short paper finds that expanding the expert selection budget (N≥K) at runtime in sparse MoE reasoning models reduces reasoning tokens by 8.5% on Qwen 35B A4B+, with no retraining or fine-tuning required.

1 source

More stories today

Open the live feed