WASTE runs Kimi K3 using 29 GB RAM at 0.50 tok/s

WASTE (Weight-Aware Streaming Tensor Engine) streams model weights to fit Kimi K3 in 29 GB of RAM, achieving 0.50 tokens/sec inference.
1 source
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- SCAIL 2 fan video replaces GTA 6 characters with fat versions
- State machine guardrails govern agent tool use in Claude Code, Cursor
- Hollywood-style racing trailer built with ComfyUI, LTX 2.3 & Krea 2
- Ethan Mollick: Microsoft and Google were daring with early AI
- Scoble: AI reads feet via video camera, no LiDAR