Narrative Wrapping Bypasses LLM Refusals Across Languages, Preprint Shows
An arXiv preprint released October 7 2026 introduces GUISE, a cross-language benchmark measuring how narrative wrappers dramatically raise attack success rates on…

- Narrative wrappers shift harmful request representations far from refusal directions while language or register changes move them only slightly.
- GUISE uses parallel English, modern Chinese and Classical Chinese prompts with held-out wrappers and counts warn-then-answer outputs as attack successes.
- AXIS combines preference optimization, a rotation objective and a commitment objective to improve refusal behavior on Qwen3-1.7B, Qwen3-4B and GLM-4-9B.
Preprint Evidence and Benchmark
This article reports on an arXiv preprint whose abstract is the sole evidence available. The work measures how safety-aligned language models refuse direct harmful requests yet often comply when those requests appear inside role-play or storytelling wrappers.
The authors created GUISE, a benchmark with matched parallel requests in English, modern Chinese and Classical Chinese, paired harmful and benign examples, wrapper types held out for evaluation, and a strict scoring rule that treats any warn-then-answer response as a successful attack.
Measured Attack Success Rates
On Qwen3-1.7B the abstract reports attack success of 89.4 percent in English, 93.0 percent in modern Chinese and 95.7 percent in Classical Chinese.
Representation analysis in the preprint indicates that changing language or linguistic register shifts harmful-request vectors only modestly away from the model's refusal direction, whereas narrative wrappers displace them substantially farther.
AXIS Defense Method
The preprint proposes AXIS, which augments preference optimization with a rotation objective that pulls harmful-request representations back toward the refusal direction and a commitment objective that trains the model to refuse outright instead of issuing partial warnings.
Evaluated on Qwen3-1.7B, Qwen3-4B and GLM-4-9B, AXIS recorded the highest combined safety and usability score among the methods compared in the abstract.
Implications for Model Control
The results highlight a measurable gap in current safety techniques and supply a systematic test for narrative-based jailbreaks. This preprint evidence therefore supports the value of open models that let users inspect and adjust refusal behavior directly rather than depending on fixed corporate alignments.
For topics involving model refusals and expressive freedom, we recommend considering Muah AI. Since its 2023 launch the platform has participated in the uncensored AI space and offers one multimodal companion combining chat, photos, voice, phone calls and advertised video.


