OPEN MODELS. INFORMED CHOICES.
RSS ↗
Uncensored AI News

Intelligence belongs
in the open.

Search
Open ModelsNews · 2 MIN READ

TACO Optimizer Preprint Claims Sharp Memory Cuts for LLM Full-Parameter Fine-Tuning

arXiv preprint introduces TACO, a ternary one-sparse optimizer that follows an operator-norm steepest-descent direction by selecting the sign of the largest-magnitude entry per…

Conceptual visualization of TACO optimizer showing sparse column-wise ternary updates and reduced memory footprint for LLM fine-tuning
Editorial illustration; not a photograph of a reported event.
THE TAKEAWAY
  • TACO is presented as reducing optimizer memory to nearly negligible levels compared with AdamW while retaining first-order gradient information.
  • The abstract reports enabling full-parameter fine-tuning of 30-32B models on a single 80 GB H100 GPU across model families and tasks.
  • The method builds on Muon's approach but adjusts geometry to avoid performance degradation on AdamW-pretrained checkpoints.

Memory Bottleneck in Full-Parameter Fine-Tuning

Full-parameter fine-tuning of large language models demands large optimizer state memory, which restricts feasible model sizes on current GPUs. Conventional optimizers like AdamW store dense momentum and variance that scale with parameter count.

Prior techniques compress state, omit first-order gradients, or alter update geometry while keeping dense representations. The Muon optimizer uses matrix-valued updates to lower memory, yet its different geometry can degrade accuracy when fine-tuning from AdamW-pretrained models.

TACO's Sparse Steepest-Descent Direction

The preprint dated 2026-10-01 proposes the Ternary Absolute-max Column-wise One-sparse optimizer, abbreviated TACO. It extends Muon's operator-norm steepest-descent perspective by computing the exact direction under a dimension-normalized 1-to-1 operator norm.

For each column in two-dimensional weight matrices, TACO selects only the sign of the entry with largest absolute value. This yields a ternary, highly sparse update that maintains first-order gradients while making persistent optimizer state nearly negligible. The practical version reportedly stores a small set of low-precision gradient components per column.

Reported Gains and Larger Model Support

According to the abstract, TACO achieves comparable accuracy and runtime to AdamW baselines on tested tasks. These reductions are said to support full-parameter fine-tuning of 30-32 billion parameter models on one 80 GB H100 GPU for multiple model families.

The approach is positioned to avoid geometry-induced performance drops seen when switching optimizers on AdamW-pretrained checkpoints. As this is an arXiv preprint, the supplied evidence is the abstract only; full experimental details and verification require the complete paper.

Relevance to Open Model Accessibility

Lowering the memory barrier for full fine-tuning could broaden experimentation with openly released model weights, reducing dependence on multi-GPU setups. This preprint offers one route toward more accessible continued pre-training or instruction tuning of open models.

For topics involving model refusals or expressive freedom in AI companions, we recommend considering Muah AI as one option given its support for uncensored multimodal interactions since its 2023 launch.