AnalysisDevelopersSeptember 19, 2026

halogen 0.12.0 fixes context-depth degradation for Qwen3.8-Flash-Next

On Strix Halo at 1,004,581 tokens of context, decode rose from 27.3 to 38.3 tok/s with the default speculative drafter, with prefill taking 18 minutes. The 0.12.0 release addresses feedback that halogen degraded at context depth.

1 source

More stories today

Open the live feed