Jailbreaking techniques pose a significant threat to the safety of Large Language Models (LLMs). Existing defenses typically focus on single-turn attacks, lack coverage across languages, and rely on limited taxonomies that either fail to capture the full diversity of attack strategies or emphasize risk categories rather than jailbreaking techniques. To advance the understanding of the effectiveness of jailbreaking techniques, we conducted a structured red-teaming challenge. The outcomes of our experiments are fourfold. First, we developed a comprehensive hierarchical taxonomy of jailbreak strategies that systematically consolidates techniques previously studied in isolation and harmonizes existing, partially overlapping classifications with explicit cross-references to prior categorizations. The taxonomy organizes jailbreak strategies into seven mechanism-oriented families: impersonation, persuasion, privilege escalation, cognitive overload, obfuscation, goal conflict, and data poisoning. Second, we analyzed the data collected from the challenge to examine the prevalence and success rates of different attack types, providing insights into how specific jailbreak strategies exploit model vulnerabilities and induce misalignment. Third, we benchmarked GPT-5 as a judge for jailbreak detection, evaluating the benefits of taxonomy-guided prompting for improving automatic detection. Finally, we compiled a new Italian dataset of 1364 multi-turn adversarial dialogues, annotated with our taxonomy, enabling the study of interactions where adversarial intent emerges gradually and succeeds in bypassing traditional safeguards.

Guarding the Guardrails: A Taxonomy-Driven Approach to Jailbreak Detection / Giarrusso, F., Sorokoletova, O., Suriani, V., Nardi, Daniele.. - 2:1(2026), pp. 191-203. (Second International Association for Safe and Ethical AI Conference (IASEAI'26) Paris; France ) [10.1609/iaseai.v2i1.43024].

Guarding the Guardrails: A Taxonomy-Driven Approach to Jailbreak Detection

Giarrusso F.
Primo
;
Sorokoletova O.
Secondo
;
Suriani V.
Penultimo
;
Nardi Daniele.
Ultimo
2026

Abstract

Jailbreaking techniques pose a significant threat to the safety of Large Language Models (LLMs). Existing defenses typically focus on single-turn attacks, lack coverage across languages, and rely on limited taxonomies that either fail to capture the full diversity of attack strategies or emphasize risk categories rather than jailbreaking techniques. To advance the understanding of the effectiveness of jailbreaking techniques, we conducted a structured red-teaming challenge. The outcomes of our experiments are fourfold. First, we developed a comprehensive hierarchical taxonomy of jailbreak strategies that systematically consolidates techniques previously studied in isolation and harmonizes existing, partially overlapping classifications with explicit cross-references to prior categorizations. The taxonomy organizes jailbreak strategies into seven mechanism-oriented families: impersonation, persuasion, privilege escalation, cognitive overload, obfuscation, goal conflict, and data poisoning. Second, we analyzed the data collected from the challenge to examine the prevalence and success rates of different attack types, providing insights into how specific jailbreak strategies exploit model vulnerabilities and induce misalignment. Third, we benchmarked GPT-5 as a judge for jailbreak detection, evaluating the benefits of taxonomy-guided prompting for improving automatic detection. Finally, we compiled a new Italian dataset of 1364 multi-turn adversarial dialogues, annotated with our taxonomy, enabling the study of interactions where adversarial intent emerges gradually and succeeds in bypassing traditional safeguards.
2026
Second International Association for Safe and Ethical AI Conference (IASEAI'26)
large language models; jailbreak detection; jailbreaking taxonomy; red teaming; multi-turn adversarial dialogues; ai safety
04 Pubblicazione in atti di convegno::04b Atto di convegno in volume
Guarding the Guardrails: A Taxonomy-Driven Approach to Jailbreak Detection / Giarrusso, F., Sorokoletova, O., Suriani, V., Nardi, Daniele.. - 2:1(2026), pp. 191-203. (Second International Association for Safe and Ethical AI Conference (IASEAI'26) Paris; France ) [10.1609/iaseai.v2i1.43024].
File allegati a questo prodotto
File Dimensione Formato  
Giarrusso_guarding-the-guardrails_2026.pdf

accesso aperto

Note: DOI: https://doi.org/10.1609/iaseai.v2i1.43024
Tipologia: Versione editoriale (versione pubblicata con il layout dell'editore)
Licenza: Tutti i diritti riservati (All rights reserved)
Dimensione 1.51 MB
Formato Adobe PDF
1.51 MB Adobe PDF

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11573/1776031
Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus ND
  • ???jsp.display-item.citation.isi??? ND
social impact