WASTE engine enables Kimi K3 inference on limited RAM

The C-based inference engine streams activated weights from NVMe, allowing the 2.78-trillion-parameter Kimi K3 model to run on 29 GB of RAM at 0.50 tokens per second. It keeps the model trunk in memory while using RAM as a bounded cache for experts.
1 source
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- Liquid AI releases LFM2.5-2.6B model for local agents
- Weaponized Email AI Assistants Could Help Attackers Hijack Accounts
- Zenity Raises $125 Million in Series C Funding
- UK regulator weighs in on whether AI scribes are medical devices
- SK hynix and SanDisk unveil High Bandwidth Flash standard for AI inference