Users share DeepSeek V4 Flash 0731 inference speeds on r/LocalLLaMA
A r/LocalLLaMA user reports ~200 tps prompt processing and ~11 tps token generation running DeepSeek V4 Flash 0731 on four 5060 Ti 16GB via llama.cpp with a 128K context window, using Unsloth's q8 lossless quant.
1 source
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- Claude Auto mode pairs with Opus 5 for multi-hour tasks
- Self-taught developer lands Director of AI role
- Humalike adds humanlike NPC behavior to Poland's biggest GTA RP server
- Sierra launches Voice Personas to give agents a brand voice and personality
- Working config shared for MiniMax Turbo LoRA (LightX2V) Ref2V audio fix