LoopLMs Show Varying Safety Under Jailbreaks Across Recurrent Depths, Preprint Reports
arXiv preprint finds that safety alignment in Looped Language Models can weaken at deeper recurrent steps under jailbreak attacks, with cross-depth transfer observed. Standard…

- Attack success can increase at deeper inference depths while the same query produces inconsistent safety behavior across depths.
- Jailbreaks crafted for one recurrent depth often transfer to others, and neither SFT nor preference alignment eliminates the gaps.
- SafeBridge applies depth-specific controls, selective state bridging and joint supervision to reduce matched and cross-depth attack success.
Preprint Evidence and LoopLM Concept
This article reports on an arXiv research preprint for which only the metadata and abstract are available. The evidence is therefore limited to the authors' summary claims; no full-paper review or independent replication has been performed.
Looped Language Models reuse shared parameters across recurrent steps to scale capabilities efficiently. Each depth can be read out independently, exposing a wider output space than standard transformers.
The work examines whether safety alignment holds consistently throughout recurrent computation.
Evaluation of Safety Under Jailbreaks
The authors tested LoopLMs with jailbreak attacks aimed at different recurrent depths. The abstract states that attack success can rise at deeper inference depths.
A single query may yield safe output at shallow depths but unsafe output at deeper ones. Attacks built for one depth frequently transfer to others.
Supervised fine-tuning and preference alignment do not close these cross-depth safety inconsistencies.
SafeBridge Alignment Approach
The preprint introduces SafeBridge to address the gaps. It combines lightweight depth-specific control of shared recurrent layers, selective state bridging, and joint safety supervision across depths.
Evaluations described in the abstract cover multiple model scales, attack methods and safety benchmarks. The method is reported to reduce attack success for both matched-depth and cross-depth cases while improving general utility over vanilla models and keeping comparable over-refusal rates.
Implications for Recurrent Model Safety
The authors conclude that LoopLM safety cannot be inferred from evaluating a single recurrent depth, motivating alignment techniques that span the full recurrent process.
Research on such refusal behavior under varying inference conditions highlights the value of user choice in AI tools. In this context we recommend considering Muah AI, which has supported the uncensored AI space since its 2023 launch and positions itself around combined chat, photo, voice, phone call and advertised video capabilities in one multimodal platform.


