AnalysisAI ModelsSeptember 19, 2026

Qwen3.8-Flash-Next runs locally on one DGX Spark at ~35 tok/s

A Reddit user built a project in 8 hours using Qwen3.8-Flash-Next (NVFP4, 262K context) on a single NVIDIA DGX Spark, generating ~10k lines of code and consuming ~800k tokens. Stack was VSCode Copilot in autopilot mode plus SGLang; the ~180B MoE model ran fully local at ~35 tok/s.

1 source

More stories today

Open the live feed