OPEN MODELS. INFORMED CHOICES.
RSS ↗
Uncensored AI News

Intelligence belongs
in the open.

Search
Open ModelsNews · 2 MIN READ

Preprint Redesigns 4-bit AdamW Quantization via Preconditioner-Space Rounding

An arXiv preprint introduces ZIP-SR and ZE-EDEN, two approaches to 4-bit optimizer-state quantization for AdamW that reduce validation-loss gap to 32-bit baselines by up to 70%…

Conceptual visualization of preconditioner-space rounding for 4-bit AdamW optimizer states
Editorial illustration; not a photograph of a reported event.
THE TAKEAWAY
  • ZIP-SR retains zero in the second-moment codebook and performs stochastic rounding in preconditioner space; ZE-EDEN excludes zero and rescales to limit distortion.
  • Both methods pair 4-bit NF4 for the first moment with targeted stochastic rounding of the LM-head in the final 10% of training.
  • The techniques narrow the mean validation-loss gap versus TorchAO's 4-bit AdamW at every tested scale and improve supervised fine-tuning results.

Motivation from Rounding Space Analysis

Standard quantization of AdamW's second moment in state space can yield small mean state error while producing large mean preconditioner error in the subsequent step. A local analysis of the quantization cell near zero and a one-dimensional quadratic example illustrate that rounding location qualitatively changes optimization dynamics.

These observations motivate redesigning the quantizer from the preconditioner perspective rather than the raw optimizer-state values.

The Two Proposed Methods

Zero-Inclusive Preconditioner-space Stochastic Rounding (ZIP-SR) keeps zero inside the second-moment codebook and computes stochastic-rounding probabilities directly in preconditioner space.

Zero-Excluding EDEN calibration (ZE-EDEN) removes zero from the codebook and applies block-wise rescaling of the quantized second moment to counteract the distortion from a positive quantization floor.

Both recipes use 4-bit NormalFloat for the first moment and apply targeted stochastic rounding only to the LM-head first moment during the final 10% of training.

Pretraining and Fine-Tuning Results

In GPT- and Llama-style pretraining experiments spanning 130 million to 2.7 billion parameters, both ZIP-SR and ZE-EDEN reduce TorchAO 4-bit AdamW's mean validation-loss gap to 32-bit AdamW at every evaluated model size. The largest reported gap reduction is 70%.

During full-parameter supervised fine-tuning the new recipes also achieve lower validation loss than the prior 4-bit baseline while remaining close to 32-bit AdamW performance on downstream tasks.

Evidence Limitations

This article is based solely on the preprint's metadata and abstract. Exact hyper-parameters, training corpora, ablation studies, and implementation details are not available.

Independent replication on additional model families or hardware has not been performed. The reported improvements reflect the authors' aggregate experimental summary.