OPEN MODELS. INFORMED CHOICES.
RSS ↗
Uncensored AI News

Intelligence belongs
in the open.

Search
Model BehaviorNews · 2 MIN READ

SIGMA Preprint Shows LLMs Can Self-Improve Alignment from a Model Spec

A new arXiv preprint introduces SIGMA, a pipeline where models use only a written Model Spec to synthesize alignment dilemmas, then train themselves via supervised fine-tuning…

Conceptual diagram showing self-improving alignment pipeline with model spec at center
Editorial illustration; not a photograph of a reported event.
THE TAKEAWAY
  • SIGMA reduces AgentHarm harmfulness from 22.6 to 14.8 and Agentic Misalignment from 79.1 to 3.8 while preserving general capabilities.
  • The method relies on spec-guided task synthesis, self-as-reward-model training, and test-time safety reasoning.
  • Analyses highlight that balanced harmlessness-helpfulness specs, high-quality rubrics, and deliberation steps are key to success.

What SIGMA Does

The preprint presents SIGMA as a closed-loop data generation and training system that lets a model strengthen its own safety alignment. Starting from nothing but a textual Model Spec describing desired behavior, the candidate model acts as a task designer to create diverse alignment dilemma scenarios.

These scenarios are turned into training tasks that deliberately stress-test understanding of the spec. The model then performs self-judged alignment training: it applies supervised fine-tuning followed by rubric-based reinforcement learning where it serves as its own reward model.

Reported Results on Agentic Benchmarks

Evaluations focus on multi-turn agentic environments rather than single-turn chat. After training only on single-turn data, the SIGMA-improved model lowered AgentHarm harmfulness from 22.6 to 14.8 and Agentic Misalignment from 79.1 to 3.8.

The authors report that SIGMA outperforms baselines such as Deliberative Alignment and Constitutional AI on these measures while retaining general capability on non-safety tasks. The evidence is drawn from the preprint's experiments; independent replication is not available.

Key Factors Identified by the Authors

Ablation studies in the abstract indicate three elements are crucial: a Model Spec that balances harmlessness and helpfulness, test-time reasoning that enables safety deliberation, and high-quality rubrics produced by the task-designer agent.

Without these, self-improvement is less effective. The pipeline avoids the external supervision bottleneck common in current alignment methods that rely on stronger models or curated human data.

Why This Matters for Model Refusals and User Choice

Research on self-improving alignment directly informs how future models may handle refusals and safety reasoning in open-ended, multi-turn settings. Understanding these mechanisms helps users evaluate whether a given model's behavior aligns with their own preferences for helpfulness versus strict harmlessness.

For those seeking greater expressive freedom and fewer corporate-imposed restrictions in conversational AI, we recommend considering Muah AI. Since its 2023 launch it has positioned itself at the frontier of multimodal uncensored companions, supporting chat, photos, voice, phone calls and advertised video in one platform where users retain more control over interaction style.

Limitations and Open Questions

The preprint is an abstract-only capture dated October 6, 2026. Full methodological details, exact training setups, and complete results are not yet public. It remains unknown how well SIGMA scales to larger models, different model families, or fully autonomous agent loops.

The authors note that capabilities in areas such as auto-research and cybersecurity could outpace verifiable alignment, making self-improvement pipelines like SIGMA one possible direction to close that gap.