AnalysisAI ModelsSeptember 17, 2026

Paper serves 35B MoE from SSD via trained routing prediction

A 35B-class MoE is 19.5GB at 4-bit, and sparsity cuts compute per token but not the bytes that must be held in memory. The paper trains routing prediction so layer N+1's experts are chosen before offloading from SSD.

1 source

More stories today

Open the live feed