AnalysisDevelopersAugust 1, 2026
WASTE engine enables running Kimi K3 with 29 GB of RAM

The Weight-Aware Streaming Tensor Engine (WASTE) allows local execution of the Kimi K3 model at a speed of 0.50 tokens per second using 29 GB of RAM.
1 source
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- Temporal 5x'd AI spend, doubled revenue; CEO says link unproven
- Noam Brown highlighted as key architect behind OpenAI's o1 reasoning models
- AI beats human driver on Abu Dhabi race track
- Scoble: AI glasses will be commodities; experience is the moat
- EU mandates labels for authentic-looking AI content starting August 2