AnalysisAI ModelsSeptember 24, 2026

Qwen 3.8 Flash Next runs 130k context on 16GB VRAM

Read original source →reddit.com

A Q4_K_M quant of Qwen 3.8 Flash Next with Q8 cache hits 15-20 t/s on a 4080 with 16GB VRAM and 64GB RAM. The setup requires a specific quant, model branch, and cache flags, documented in a GitHub repo.

1 source

More stories today

Open the live feed