AnalysisAI ModelsAugust 28, 2026

Audit finds 64 of 443 GGUF quants mislabeled

An audit of 443 GGUF quants across 25 repos found 64 files whose filenames claim a lower bit-width than they actually are, due to llama-quantize silently swapping in a ~4.5 bpw type when tensor rows aren't divisible by 256. On Nemotron-3.5-Lightning, all four IQ2 rungs are the same 4.58 bpw file.

1 source

Daily brief

Get tomorrow's AI brief in your inbox

More stories today

Open the live feed