Responsible AI initiatives place great emphasis on the safety of Large Language Model (LLM)-based systems. In particular, it has become standard practice to subject these models to an alignment procedure aimed at preventing harmful outputs. However, once aligned, a model is not guaranteed to maintain this alignment throughout its lifecycle. Moreover, the likelihood of misalignment increases as malicious actors may deliberately employ jailbreaking techniques to compromise LLM safety. To counter this, much research has focused on improving alignment methods and post-processing filters. In this paper, we introduce a new perspective on advancing LLM alignment: rather than developing stronger alignment techniques, we investigate the model’s intrinsic ability to recover its alignment after corruption. We propose a methodology for modeling the safety trajectories of user-assistant interactions and for detecting recovery trends within them. We apply this approach to a jailbreaking scenario, presenting a preliminary recovery analysis based on a dataset of adversarial multi-turn dialogues and examining the influence of the content moderation model chosen for safety evaluation. Project page with an interactive data visualizer is available at https://lab-rococo-sapienza.github.io/LearningfromMistakes/.

Learning from Mistakes: Can LLM Self-Recover after Misalignment? / Sorokoletova, O., Giarrusso, F., Suriani, V., Nardi, D.. - 4189:(2026), pp. 74-83. (2026 Machine Ethics: From Formal Methods to Emergent Machine Ethics Workshop, MEW 2026 Singapore ).

Learning from Mistakes: Can LLM Self-Recover after Misalignment?

Sorokoletova O.
Primo
;
Giarrusso F.
Secondo
;
Suriani V.
Penultimo
;
Nardi D.
Ultimo
2026

Abstract

Responsible AI initiatives place great emphasis on the safety of Large Language Model (LLM)-based systems. In particular, it has become standard practice to subject these models to an alignment procedure aimed at preventing harmful outputs. However, once aligned, a model is not guaranteed to maintain this alignment throughout its lifecycle. Moreover, the likelihood of misalignment increases as malicious actors may deliberately employ jailbreaking techniques to compromise LLM safety. To counter this, much research has focused on improving alignment methods and post-processing filters. In this paper, we introduce a new perspective on advancing LLM alignment: rather than developing stronger alignment techniques, we investigate the model’s intrinsic ability to recover its alignment after corruption. We propose a methodology for modeling the safety trajectories of user-assistant interactions and for detecting recovery trends within them. We apply this approach to a jailbreaking scenario, presenting a preliminary recovery analysis based on a dataset of adversarial multi-turn dialogues and examining the influence of the content moderation model chosen for safety evaluation. Project page with an interactive data visualizer is available at https://lab-rococo-sapienza.github.io/LearningfromMistakes/.
2026
2026 Machine Ethics: From Formal Methods to Emergent Machine Ethics Workshop, MEW 2026
Large Language Model (LLM), LLM Safety, Alignment, Safeguarding, Content Moderation Tools, Jailbreaking
04 Pubblicazione in atti di convegno::04b Atto di convegno in volume
Learning from Mistakes: Can LLM Self-Recover after Misalignment? / Sorokoletova, O., Giarrusso, F., Suriani, V., Nardi, D.. - 4189:(2026), pp. 74-83. (2026 Machine Ethics: From Formal Methods to Emergent Machine Ethics Workshop, MEW 2026 Singapore ).
File allegati a questo prodotto
File Dimensione Formato  
Sorokoletova_Learning_2026.pdf

accesso aperto

Note: https://ceur-ws.org/Vol-4189/paper5.pdf
Tipologia: Versione editoriale (versione pubblicata con il layout dell'editore)
Licenza: Creative commons
Dimensione 1.41 MB
Formato Adobe PDF
1.41 MB Adobe PDF

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11573/1776121
Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus 0
  • ???jsp.display-item.citation.isi??? ND
social impact