LaunchDevelopersAugust 1, 2026

WASTE runs Kimi K3 using 29 GB RAM at 0.50 tok/s

WASTE (Weight-Aware Streaming Tensor Engine) streams model weights to fit Kimi K3 in 29 GB of RAM, achieving 0.50 tokens/sec inference.

1 source

Daily brief

Get tomorrow's AI brief in your inbox

More stories today

Open the live feed