AnalysisAI ModelsAugust 30, 2026

EXL3 quants run dense 30B model on 12GB VRAM

A user runs Muse Glimmer 30B EXL3-SC 3.00bpw H4 fully on a 12GB VRAM GPU at 100K context with Q8_0 KV cache, achieving ~30 tok/s. Claims it's only slightly worse than the official 17GB quant.

1 source

More stories today

Open the live feed