AnalysisAI ModelsJuly 29, 2026

Kimi K3 runs at ~4 tokens/s on home lab via llama.cpp fork

Reddit user iVoider reports ~4 tokens/s running Kimi K3 locally on 768GB DDR5 with two RTX 5090s, using GrEarl's Q2_K GGUF and pwilkin's llama.cpp 'kimi-k3-text' fork — calling the results 'better than expected'.

1 source

Daily brief

Get tomorrow's AI brief in your inbox

More stories today

Open the live feed