AnalysisAI ModelsAugust 30, 2026

Qwen3.8-Flash-Next 2-bit quant runs 350K context on 128 GB M5 Max

A 78.9 GB UD-Q2_K_XL Unsloth quant of Qwen3.8-Flash-Next ran a 358,400-token context slot on a 128 GB M5 Max MacBook Pro via llama.cpp b10686, using YaRN to extend from the native 262,144 tokens. The 3.5-hour, 100-turn test charted speed against context depth.

1 source

More stories today

Open the live feed