Lower Reconstruction Loss Can Degrade Low-Bit LLM Quantization, DRQ Refinement Proposed
arXiv preprint finds that minimizing reconstruction loss in weight-only PTQ does not guarantee better model performance and can hurt results when input activation distributions…

- Weights selected by lower reconstruction loss on calibration data can produce worse performance on the same data or new tasks.
- DRQ minimizes worst-case reconstruction loss over varied input activation distributions without changing quantization parameters or inference operators.
- The method refines integer codes within the existing grid and improves quantized models according to the reported experiments.
Core Observation from the Abstract
The preprint shows that minimizing reconstruction loss in weight-only post-training quantization does not always select weights that improve model performance on new tasks. Lower reconstruction loss can even degrade performance on the calibration data itself.
Analysis indicates that weights favored under calibration inputs may incur higher loss when the distribution of input activations changes. This observation motivates moving beyond average-case optimization.
Distributionally Robust Quantization
The authors propose DRQ, a post-hoc refinement that minimizes the worst-case reconstruction loss over a constrained set of possible input activation distributions.
DRQ adjusts only the integer codes of quantized weights inside the existing quantization grid. It leaves scaling parameters and inference operators unchanged, adding no extra cost at inference time.
Reported Experimental Outcomes
According to the abstract, DRQ improves downstream performance of models quantized by six representative PTQ methods, including AWQ, GPTQ, and ParoQuant. Gains appear on both dense and mixture-of-experts large language models.
The work positions DRQ as a general post-hoc framework for weight-only PTQ that enhances results without inference overhead. This evidence is drawn from the preprint abstract only.
Evidence Limitations
Only the arXiv abstract is available in this capture. Full details on methods, proofs, and exact experimental setups remain unknown until the complete paper is released.
Claims in the abstract should be treated as preliminary. Independent replication or deeper analysis must await the full preprint.


