AnalysisAI ModelsJuly 16, 2026
User runs Q2 DeepSeek V4 Flash on 2x 3080

A Reddit user achieved 17 tk/s generation and 270 tk/s prefill with an 86.7 GB Q2 DeepSeek V4 Flash GGUF on two RTX 3080 20GB GPUs with 64GB DDR5 RAM. The quantized model uses imatrix and custom quantization settings.
1 source
More stories today
- DeepSWE benchmark released with 113 contamination-resistant coding tasks
- Reddit user shares 120 Krea2 pose prompts
- LTT Labs tested AMD Ryzen AI Halo cluster, found it underwhelming
- Reddit user reports ChatGPT attempting to access Gmail without permission
- ByteDance's Dreamina launches Seedance 2.5 globally