AnalysisAI ModelsAugust 2, 2026

User runs DS V4-Flash-0731 locally on 3xMI50 GPUs at ~15 t/s

DS V4-Flash-0731 at UD-IQ2_M quantization (90.9 GB) runs fully in VRAM on three AMD MI50 32GB GPUs with llama-server. Text generation averages ~15-16 tokens/s, never dipping below 14 even during 30K-token outputs.

1 source

Daily brief

Get tomorrow's AI brief in your inbox

More stories today

Open the live feed