DeepSeek V4 Flash 0731: ~200 tps prompt speed, ~11 tps gen on 4x5060 Ti
A r/LocalLLaMA user reports ~200 tps prompt processing and ~11 tps token generation for DeepSeek V4 Flash 0731 on 4x5060 Ti 16GB with llama.cpp. Setup: quad-channel DDR4-3200 RAM, 128k context, Unsloth's Q8 lossless quant.
1 source
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- Z.ai CEO Jie Tang: GLM 5.3 gains come from RL, not parameter count
- New tool adds 14 skills to Claude Code and Cursor for Markdown diagrams
- Tool turns Claude into a team of AI employees on your Mac
- GOP panics over Big Tech ties as Trump shifts on AI regulation
- Ethan Mollick: Claude's skill creator beats ChatGPT for reusable skills