DeepSeek-V4-Flash-0731 UD-Q8_K_XL hits 17.20 t/s on RTX A6000
Local inference held a steady 17.20 t/s on an RTX A6000 48GB with 256GB DDR4; prompt processing ran in the high 70 t/s. The 48GB VRAM fit the full 1M-token context, dropping to mid-30 t/s at 300k context.
1 source
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- Meta AI desktop app for macOS now available to download
- OpenAI model solves ten open problems in math and computer science
- Tacta Systems launches TactaBot robotic hand for skilled manufacturing
- Rippling launched AI Spend Console after burning millions on AI tokens
- Claude Code sessions can now talk to each other