ResidualQuant Preprint Proposes 2-Bit KV Cache Quantization for Looped Transformers
An arXiv preprint dated 2026-10-07 introduces ResidualQuant, a KV cache quantization technique for looped Transformers. It leverages similarity of KV states across loops to…

- ResidualQuant treats final-loop KV states as reference and quantizes earlier loops with low-precision residuals, adding least-squares scaling, rotations, and loop-wise mixed precision.
- The abstract claims the method improves accuracy-memory tradeoff over rotation-based KV quantization on mathematical reasoning and code generation benchmarks.
- It reports theoretical KV storage reduction of 80.7% while retaining accuracy close to BF16 in mixed-precision settings, with hardware gains on an RTX 5090.
Looped Transformers and the Memory Bottleneck
Looped Transformers reuse shared blocks across recurrent loops to increase computational depth while keeping parameter count fixed. This design improves efficiency but causes KV cache memory to scale with the number of loops.
The growing cache limits batch sizes and inference throughput. Although KV cache quantization can reduce memory use, prior methods often degrade accuracy at low bit widths such as 2 bits.
How ResidualQuant Works
The preprint observes that KV states across loops in these models are highly similar. ResidualQuant therefore uses the final-loop KV states as a high-precision reference and represents earlier loops with low-precision residuals.
The technique further applies least-square scaling and rotations to the residuals along with loop-wise mixed precision. This supports accurate reconstruction at INT2 precision.
Reported Performance Claims
According to the abstract, ResidualQuant delivers a better accuracy-memory tradeoff than state-of-the-art rotation-based KV quantization across looped Transformer models and benchmarks focused on mathematical reasoning and code generation.
The authors state that under mixed-precision settings the approach retains accuracy close to BF16 while cutting theoretical KV storage by 80.7 percent. It reportedly achieves up to 13.0 percent higher accuracy than the baseline at the same memory budget.
Hardware Results on RTX 5090
The preprint reports that reduced KV memory traffic improves fixed-batch decode throughput by up to 2.73 times on an RTX 5090. The smaller memory footprint also enables up to twice larger batches, boosting peak throughput by as much as 4.15 times.
Evidence Limitations and Context
This summary is based solely on the arXiv abstract and metadata captured on 2026-10-07. The full paper may contain additional experimental details, proofs, or implementation specifics that are not available here.
Quantization methods like ResidualQuant can help reduce memory demands for running larger or deeper models locally. Such advances support user choice in open models and local inference where individuals control their AI tools and data.


