AnalysisAI ModelsJuly 29, 2026
Kimi K3 runs at ~4t/s on home lab with 768GB DDR5 and 2x5090

A Reddit user reports running Kimi K3 on a home setup with 768GB DDR5 and 2x5090 GPUs, achieving approximately 4 tokens per second using a custom llama.cpp fork and GGUF quantized weights.
1 source
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- PortSwigger explains safety design for agentic pentesting
- Kimi K3 distillation into Laguna 2.1 requested
- Cursor and Anthropic launch localized India pricing plans
- SpaceXAI releases Grok Voice Think Fast 2.0
- Sam Altman to brief White House on OpenAI's next AI model