DeepSeek-V4-Flash-0731 UD-Q8_K_XL hits 17.20 t/s on A6000
User runs DeepSeek-V4-Flash-0731 UD-Q8_K_XL on AMD EPYC 74F3, RTX A6000 48GB and 8-channel DDR4. Inference is steady 17.20 t/s; 48GB VRAM fits the full 1M context, with prompt processing around 70 t/s.
1 source
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- Z.ai CEO Jie Tang: GLM 5.3 gains come from RL, not parameter count
- New tool adds 14 skills to Claude Code and Cursor for Markdown diagrams
- Tool turns Claude into a team of AI employees on your Mac
- GOP panics over Big Tech ties as Trump shifts on AI regulation
- Ethan Mollick: Claude's skill creator beats ChatGPT for reusable skills