This report presents the contribution from the DIAG-Sapienza team for the GSI:detect task at EVALITA 2026. We address both the main regression task and the classification sub-task in a joint formulation, motivated by the hypothesis that category-level information can support more accurate estimation of gender stereotype intensity, especially in low-data settings. Our contribution features three systems corresponding to the challenge tracks: 1) a zero-shot approach that used GPT-5 as an LLM-as-a-Judge with structured prompting and explanation generation to elicit implicit reasoning, 2) a few-shot approach that augments prompting with randomly sampled and retrieved in-context examples, and 3) an encoder-only fine-tuned RoBERTa-based model that integrates LLM-generated reasoning sentences as additional input. Leaderboard results on the official test set show that our zero-shot and few-shot systems rank first and second, on the main regression task, outperforming all competing submissions. Overall, the results suggest that jointly predicting gender stereotype values and categories benefits the regression task, and that eliciting concise explanations can improve prediction quality.
DIAG-Sapienza at GSI:detect: Joint Detection and Classification of Gender Stereotypes with Structured Prompting and Fine-Tuning / Sorokoletova, O., Musumeci, E., Nardi, D.. - 4195:(2026). (EVALITA 2026 9th Evaluation Campaign of Natural Language Processing and Speech Tools for Italian Bari, Italy ).
DIAG-Sapienza at GSI:detect: Joint Detection and Classification of Gender Stereotypes with Structured Prompting and Fine-Tuning
Sorokoletova, Olga
;Musumeci, Emanuele
;Nardi, Daniele
2026
Abstract
This report presents the contribution from the DIAG-Sapienza team for the GSI:detect task at EVALITA 2026. We address both the main regression task and the classification sub-task in a joint formulation, motivated by the hypothesis that category-level information can support more accurate estimation of gender stereotype intensity, especially in low-data settings. Our contribution features three systems corresponding to the challenge tracks: 1) a zero-shot approach that used GPT-5 as an LLM-as-a-Judge with structured prompting and explanation generation to elicit implicit reasoning, 2) a few-shot approach that augments prompting with randomly sampled and retrieved in-context examples, and 3) an encoder-only fine-tuned RoBERTa-based model that integrates LLM-generated reasoning sentences as additional input. Leaderboard results on the official test set show that our zero-shot and few-shot systems rank first and second, on the main regression task, outperforming all competing submissions. Overall, the results suggest that jointly predicting gender stereotype values and categories benefits the regression task, and that eliciting concise explanations can improve prediction quality.| File | Dimensione | Formato | |
|---|---|---|---|
|
Sorokoletova_DIAG-Sapienza_2026.pdf
accesso aperto
Note: https://ceur-ws.org/Vol-4195/43.pdf
Tipologia:
Versione editoriale (versione pubblicata con il layout dell'editore)
Licenza:
Creative commons
Dimensione
494.57 kB
Formato
Adobe PDF
|
494.57 kB | Adobe PDF |
I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.


