How-ToAI ModelsSeptember 14, 2026

Qwen3.8-Flash-Next runs on a 12GB VRAM card at 20.65 tok/s

Read original source →carteakey.dev

A 125B-A6B MoE with a 51B n-gram lookup table, quantized to 88 GiB, decodes at 19.35-20.65 tok/s on an RTX 4070 12GB + 64GB DDR5. The n-gram table is hashed 3-token lookups, so it never needs to sit in RAM.

1 source

More stories today

Open the live feed