A team led by Zhuo Sun released a method called Distributionally Robust Quantization (DRQ) on arXiv. DRQ adjusts the integer codes of quantized weights to minimise worst‑case reconstruction loss across varied activation distributions. It works after weight‑only post‑training quantization and leaves the quantization grid and inference operators unchanged. Experiments show DRQ improves models already quantised by six methods, such as AWQ, GPTQ, and ParoQuant, for both dense and mixture‑of‑experts LLMs.
Why it matters
It lets developers deploy smaller, faster LLMs with better accuracy without changing hardware or adding runtime overhead.