WASTE runs full 2.78T-parameter Kimi K3 by streaming experts from NVMe

WASTE is a dependency-free, embeddable C inference engine. It keeps the model trunk in memory, streams activated experts directly from disk, and uses remaining RAM as a bounded expert cache to run Kimi K3 beyond available memory.
1 source
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- Law firms urged to reclaim AI sovereignty from frontier labs
- Thread ping-back graph as primitive multiagent AGI future
- OpenAI asks judge to toss Apple's trade secret lawsuit
- Five-parallel-agent AI sales team runs inside Claude Code
- Deep Eye AI pen-testing tool scans 45+ vulnerability types