LaunchDevelopersAugust 3, 2026

WASTE inference engine streams Kimi K3's 2.78T parameters from NVMe

The dependency-free C engine keeps the model trunk in memory, streams activated experts from disk, and uses remaining RAM as a bounded expert cache.

1 source

Daily brief

Get tomorrow's AI brief in your inbox

More stories today

Open the live feed