AnalysisDevelopersSeptember 11, 2026

Cherenkov runs Qwen3.8-Flash-Next on 32GB M4 MacBook Air at 8-22 tok/s

Cherenkov, an Apple Silicon inference engine using predictive expert streaming, runs Qwen3.8-Flash-Next in Q4 with only ~21GB of allocations. The author claims a memory-constrained inference record for the model on Apple Silicon.

1 source

More stories today

Open the live feed