AnalysisAI ModelsSeptember 30, 2026

Qwen3.8 Flash Next GGUF hits 50t/s on 12GB VRAM laptop via Strata engine

Read original source →reddit.com

Reddit user reports 50 tokens/s generation and 1500 tokens/s prompt processing running the ISTA-DASLab GGUF of Qwen3.8 Flash Next on a 12GB VRAM, 64GB RAM laptop. The Strata inference engine currently runs only on Nvidia GPUs.

1 source

More stories today

Open the live feed