Quantization's hidden cost: multilingual and decision-level damage
New arXiv papers show low-bit quantization disproportionately harms non-English languages (2-4x perplexity degradation in sub-4B INT3 GPTQ) and can flip MoE routing decisions. Proposals include Language-Conditional Dequantization and methods to predict which decisions break.
How this story unfolded
3 days · 14 reports · 14 of 15 shown
- Aug 10
EntropyMoE: Entropy-Aware Sparse Expert Routing for Tokenizer-Free LLMsarxiv.org
TEXAS: Task-Expert-Aware Supervision for Downstream Mixture-of-Experts LLM Adaptationarxiv.org
Beyond Routing Weights: Faithful Response-Level Interpretation of Mixture-of-Experts Reward Models via Contribution Contrastarxiv.org
Quantization Damage Is Multiplicative, Not Additivearxiv.org
- Aug 11
Shape Mutating Expert Compression:LorExperts and BTExpertsarxiv.org
Router Sensitivity Under Lightweight Fine-Tuning Identifies Prunable Experts in Mixture-of-Experts Modelsarxiv.org
When Does Trace-Driven Evaluation Mislead MoE Expert Caching? Replay Semantics, Workload Contamination, and Operating Regimesarxiv.org
EasyBalance: Cross-Layer Load Balancing in Distributed MoE Inferencearxiv.org
The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelismarxiv.org
- Aug 12
- Aug 13
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- Z.ai CEO Jie Tang: GLM 5.3 gains come from RL, not parameter count
- New tool adds 14 skills to Claude Code and Cursor for Markdown diagrams
- Tool turns Claude into a team of AI employees on your Mac
- GOP panics over Big Tech ties as Trump shifts on AI regulation
- Ethan Mollick: Claude's skill creator beats ChatGPT for reusable skills