WASTE inference engine enables running Kimi K3 model from NVMe

The C-based engine allows running the 2.78-trillion-parameter Kimi K3 model by streaming activated weights from NVMe storage, using RAM as a cache for experts. It is dependency-free and designed to be embeddable.
1 source
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- Meta AI desktop app for macOS now available to download
- OpenAI model solves ten open problems in math and computer science
- Tacta Systems launches TactaBot robotic hand for skilled manufacturing
- Rippling launched AI Spend Console after burning millions on AI tokens
- Claude Code sessions can now talk to each other