DeepSeek V4 Flash hits ~3.5 tok/s at IQ2_M on dual RTX 3060

DeepSeek V4 Flash 0731 with IQ2_M quantization averaged ~3.5 tok/s on dual RTX 3060 GPUs with 96GB RAM. The user switched to Unsloth Studio after LM Studio refused to load weights onto the second GPU, and cautioned the result was not a proper benchmark.
1 source
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- Claude Auto mode pairs with Opus 5 for multi-hour tasks
- Self-taught developer lands Director of AI role
- Humalike adds humanlike NPC behavior to Poland's biggest GTA RP server
- Sierra launches Voice Personas to give agents a brand voice and personality
- Working config shared for MiniMax Turbo LoRA (LightX2V) Ref2V audio fix