DeepSeek-V4-Flash-0731 hits 12.5 tok/s on RTX 3090 + 128GB DDR5

User runs DeepSeek-V4-Flash-0731 (UD-IQ3_S quantization) at 12.5 tok/s on a 24GB RTX 3090 with 128GB DDR5 at 5600 MHz, using text-generation-webui's llama.cpp loader. Setup required replacing the bundled llama.cpp binaries as a workaround.
1 source
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- Meta AI desktop app for macOS now available to download
- OpenAI model solves ten open problems in math and computer science
- Tacta Systems launches TactaBot robotic hand for skilled manufacturing
- Rippling launched AI Spend Console after burning millions on AI tokens
- Claude Code sessions can now talk to each other