User runs DS V4-Flash-0731 locally on 3xMI50 GPUs at ~15 t/s

DS V4-Flash-0731 at UD-IQ2_M quantization (90.9 GB) runs fully in VRAM on three AMD MI50 32GB GPUs with llama-server. Text generation averages ~15-16 tokens/s, never dipping below 14 even during 30K-token outputs.
1 source
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- Meta AI desktop app for macOS now available to download
- OpenAI model solves ten open problems in math and computer science
- Tacta Systems launches TactaBot robotic hand for skilled manufacturing
- Rippling launched AI Spend Console after burning millions on AI tokens
- Claude Code sessions can now talk to each other