WASTE inference engine enables running Kimi K3 model from NVMe

WASTE is a dependency-free C inference engine that runs the 2.78-trillion-parameter Kimi K3 model by streaming activated weights from disk. It maintains the model trunk in RAM while using remaining memory as a cache for experts.
1 source
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- Legal tech sector sees wave of AI startup acquisitions
- Embodied-AI data startup Kaiwang Data raises RMB100M+
- rust-lang/rust is adopting an LLM policy
- ByteDance launches SeedRealtime full-duplex audio-video model
- AirLLM loads layers to run 2.8T-parameter Kimi K3 on 4GB GPU