AnalysisAI ModelsJuly 16, 2026

User runs Q2 DeepSeek V4 Flash on 2x 3080

A Reddit user achieved 17 tk/s generation and 270 tk/s prefill with an 86.7 GB Q2 DeepSeek V4 Flash GGUF on two RTX 3080 20GB GPUs with 64GB DDR5 RAM. The quantized model uses imatrix and custom quantization settings.

1 source

More stories today

Open the live feed