Align-Then-Correct: Training-Free Two-Stage Compensation for 2-Bit LLMs
arXiv preprint abstract introduces Align-then-Correct, a training-free low-rank quantization error compensation method that addresses symmetric calibration and second-order…

- The framework uses two closed-form stages: first aligning layer outputs to the full-precision model via a Fisher-weighted asymmetric objective, then applying a rank-constrained natural-gradient step after re-measuring statistics.
- Each low-rank adapter is computed with a single truncated SVD; backward passes collect statistics only.
- At 2-bit quantization under QuIP#, the method improves WikiText-2 perplexity and recovers more of the accuracy gap to FP16 on C4 versus baselines, with gains also reported on zero-shot tasks.
Limitations Identified in Existing LQEC Methods
Low-rank quantization error compensation adds a closed-form rank-r adapter beside each frozen quantized weight to recover accuracy without training. The preprint abstract notes that current compensators share two simplifications limiting their effectiveness.
Symmetric calibration uses the same activations for full-precision and compensated weights, resulting in a high-rank compensation target where a fixed rank budget recovers only a fraction of the error. Additionally, methods minimize only the second-order loss term despite a remaining first-order descent direction in every layer of the non-stationary compensated model.
Proposed Two-Stage Framework
Align-then-Correct removes both simplifications through a two-stage closed-form process. Stage 1 aligns each layer's output to the full-precision model using a Fisher-weighted asymmetric objective, producing a lower-rank target that better utilizes the rank budget.
Stage 2 re-measures statistics on the compensated model and applies a rank-constrained natural-gradient step to absorb the leftover first-order signal. All adapters are derived from a single truncated SVD per layer.
Results from the Preprint Abstract
Evaluated at 2 bits under QuIP#, the approach reduces WikiText-2 perplexity from 12.43 to 10.26 on Qwen3-8B and from 21.11 to 13.22 on Qwen3-4B. On held-out C4 it recovers 51% and 84% of the gap to FP16, compared with 31% and 63% for the strongest baseline.
Consistent improvements are noted in seven-task zero-shot average accuracy, at higher bit-widths, and under a different quantizer. This evidence is limited to the arXiv abstract dated October 6, 2026; full methodological details and results require the complete paper.


