AnalysisAI ModelsSeptember 30, 2026

Qwen3.8 Flash Next GGUF hits 50t/s on 12GB VRAM laptop via Strata engine

Read original source →reddit.com

Reddit user reports the ISTA-DASLab GGUF of Qwen3.8 Flash Next running at 50 tokens/s generation and 1500 tokens/s prompt processing on a 12GB VRAM, 64GB RAM laptop using the Strata inference engine. Strata currently runs only on Nvidia GPUs; the developer fixed early kv-cache and CPU-throttling bugs.

1 source

More stories today

Open the live feed