Kimi K3 runs at ~4 tokens/s on home lab via llama.cpp fork

Reddit user iVoider reports ~4 tokens/s running Kimi K3 locally on 768GB DDR5 with two RTX 5090s, using GrEarl's Q2_K GGUF and pwilkin's llama.cpp 'kimi-k3-text' fork — calling the results 'better than expected'.
1 source
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- Orchestrate Claude Code as a multi-agent startup team
- Walkthrough shows building production-ready systems with AI-assisted development
- Claude Code tool generates 23 types of Mermaid diagrams
- Developer adds LoRA motion training support for MiniMax H3
- Meta appears to be expanding its web index for Meta AI