DeepSeek-V4-Flash-0731 UD-IQ3_S: 12.5 tok/s on RTX 3090

Run via llama.cpp in text-generation-webui on an RTX 3090 24GB with 128GB DDR5 at 5600 MHz (AMD EXPO). The user had to replace text-generation-webui's bundled llama.cpp binaries as a workaround.
1 source
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- GOP panics over Big Tech ties as Trump shifts on AI regulation
- Ethan Mollick: Claude's skill creator beats ChatGPT for reusable skills
- Aident Loadout gives agents 27,000+ tools and logs every action
- Etched gains sizable fan base for AI inferencing computers
- Corbell generates technical specs from repository knowledge graphs