WASTE runs full 2.78T-parameter Kimi K3 on 29 GB RAM via NVMe streaming

A dependency-free, embeddable C inference engine, WASTE runs the full 2.78-trillion-parameter Kimi K3 using just 29 GB of RAM at 0.50 tok/s. It keeps the model trunk in memory, streams selected experts directly from NVMe, and uses remaining RAM as a bounded expert cache.
1 source
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- CoreWeave to Enter Asian Market With Indonesian Data Centers
- Nvidia, Dell Back AI Cloud Startup Volta at $2.4 Billion Value
- ESPN unveils AI tells detection at World Series of Poker
- Gemini Agent-to-Agent Attack Exposed Secrets, Enabled Pull Request Tampering
- Podcast examines Hollywood's quiet embrace of AI and control battle