LaunchDevelopersAugust 3, 2026

WASTE runs full 2.78T-parameter Kimi K3 by streaming experts from NVMe

WASTE is a dependency-free, embeddable C inference engine. It keeps the model trunk in memory, streams activated experts directly from disk, and uses remaining RAM as a bounded expert cache to run Kimi K3 beyond available memory.

1 source

Daily brief

Get tomorrow's AI brief in your inbox

More stories today

Open the live feed