LaunchDevelopersJuly 31, 2026

WASTE runs full 2.78T-parameter Kimi K3 on 29 GB RAM via NVMe streaming

A dependency-free, embeddable C inference engine, WASTE runs the full 2.78-trillion-parameter Kimi K3 using just 29 GB of RAM at 0.50 tok/s. It keeps the model trunk in memory, streams selected experts directly from NVMe, and uses remaining RAM as a bounded expert cache.

1 source

Daily brief

Get tomorrow's AI brief in your inbox

More stories today

Open the live feed