AnalysisAI ModelsSeptember 19, 2026

Qwen3.8-Flash-Next runs locally on one DGX Spark at ~35 tok/s

Read original source →reddit.com

A Reddit user reports generating ~10k lines of code in 8 hours with Qwen3.8-Flash-Next (NVFP4, 262K context) on a single NVIDIA DGX Spark, consuming ~800k tokens. Stack was VSCode Copilot autopilot plus SGLang; the ~180B MoE ran fully local at ~35 tok/s.

1 source

More stories today

Open the live feed