AnalysisAI ModelsAugust 30, 2026

EXL3 quants run Muse Glimmer 30B on 12GB VRAM

A Reddit user reports running Muse Glimmer 30B EXL3-SC 3.00bpw H4 fully on a 12GB VRAM GPU at 100K context with Q8_O KV cache, achieving ~30 tok/s on a VRAM-constrained laptop. The quant is described as only slightly worse than the official 17GB version.

1 source

More stories today

Open the live feed