AnalysisAI ModelsSeptember 16, 2026

Qwen3.8 Flash runs on 12GB VRAM at 15 tokens/s

Read original source →reddit.com

A Reddit user reports steady 15 tokens/s output and 100-120 tokens/s prompt processing on an RTX 5070 with 12GB VRAM using a Qwen3.8-Flash GGUF at 3 bpw (IQ3_XXS). The full GGUF is ~76GB, with only 47GB sharded into VRAM and RAM.

1 source

More stories today

Open the live feed