New research advances LLM post-training quantization techniques

Recent papers introduce ReQuant and CubicQuant to improve model compression, while researchers identify that quantization damage is multiplicative rather than additive. A community project, KLQ, also released a training-free rotation-based quantization method for Llama 3.2 1B.
How this story unfolded
1 day · 3 reports · 1 community post · 4 of 5 shown
- Aug 10
CubicQuant: Parametric Non-Uniform Codebooks for High-Throughput LLM Inference with 1-8-Bit Weightsarxiv.org
ReQuant: Fixed-Grid Discrete Refinement for Post-Training Quantizationarxiv.org
KLQ: Training-free measured rotation quantization. Beats all training-free rotation-based quantization methods on W4A4KV4-bits. Llama 3.2 1B KLQ-quantized beats SpinQuant and gets close to ReSpinQuant without GPTQ/LDLQ rounding.
- Aug 11
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- VisionDepth3D creates 3D from 2D videos with AI depth mapping
- xAI SDK v1.18.0 adds grok-4.6 and xhigh reasoning_effort
- WhisperKit enables local speech-to-text transcription on macOS
- Autoware accelerates autonomous vehicle deployment with open-source stack
- Perplexity reportedly offered to acquire Google Chrome one year ago