AnalysisAI ModelsAugust 2, 2026

DeepSeek-V4-Flash-0731 UD-Q8_K_XL hits 17.20 t/s on A6000

Read original source →reddit.com

User runs DeepSeek-V4-Flash-0731 UD-Q8_K_XL on AMD EPYC 74F3, RTX A6000 48GB and 8-channel DDR4. Inference is steady 17.20 t/s; 48GB VRAM fits the full 1M context, with prompt processing around 70 t/s.

1 source

More stories today

Open the live feed