AnalysisAI ModelsSeptember 30, 2026

Qwen3.8 Flash Next GGUF hits 50t/s on 12GB VRAM laptop via Strata engine

Read original source →reddit.com

A Reddit user reports running the ISTA-DASLab GGUF of Qwen3.8 Flash Next at 50 tokens/s generation and 1500 tokens/s prompt processing on a 12GB VRAM, 64GB RAM laptop using the Strata inference engine. Strata currently runs only on Nvidia GPUs.

1 source

More stories today

Open the live feed