Personal data are routinely de-identified before release. However, re-identification may still occur through linkage with external information. This raises the problem of evaluating the risk of re-identification in supposedly anonymous data. In this work, we propose an approach called CIRE (Classification-based Identification Risk Evaluation) to quantify such risk by combining statistical analysis of the dataset with a classification-based method designed to detect low-density regions of the data space. Individuals falling within these sparse regions face a particularly high re-identification risk, since unique or rare combinations of features, even after de-identification, may inadvertently single them out. The proposed technique is evaluated on a large real-world dataset of 22,758,797 records and 37 variables from the Italian National Statistical Office. After a training phase requiring 1360 s, CIRE is able to evaluate the risk of each million of new records in about 39 s. The fraction of the cases classified as high-risk in the full dataset is on the order of (Formula presented), with each of these records receiving an explicit risk assessment. When compared with Isolation Forest, CIRE more accurately finds localized sparse and even previously unobserved regions that are critical for the assessment of identification risk. These results demonstrate that CIRE is a scalable, interpretable, and operational tool for real-time evaluation of re-identification risk.

A classification-based approach to the evaluation of personal identification risk for supposedly anonymous data / Bruni, R., Bianchi, G., Lorusso, P.. - In: INFORMATION PROCESSING & MANAGEMENT. - ISSN 1873-5371. - 64:1(2026). [10.1016/j.ipm.2026.105073]

A classification-based approach to the evaluation of personal identification risk for supposedly anonymous data

Renato Bruni
Primo
;
Gianpiero Bianchi;
2026

Abstract

Personal data are routinely de-identified before release. However, re-identification may still occur through linkage with external information. This raises the problem of evaluating the risk of re-identification in supposedly anonymous data. In this work, we propose an approach called CIRE (Classification-based Identification Risk Evaluation) to quantify such risk by combining statistical analysis of the dataset with a classification-based method designed to detect low-density regions of the data space. Individuals falling within these sparse regions face a particularly high re-identification risk, since unique or rare combinations of features, even after de-identification, may inadvertently single them out. The proposed technique is evaluated on a large real-world dataset of 22,758,797 records and 37 variables from the Italian National Statistical Office. After a training phase requiring 1360 s, CIRE is able to evaluate the risk of each million of new records in about 39 s. The fraction of the cases classified as high-risk in the full dataset is on the order of (Formula presented), with each of these records receiving an explicit risk assessment. When compared with Isolation Forest, CIRE more accurately finds localized sparse and even previously unobserved regions that are critical for the assessment of identification risk. These results demonstrate that CIRE is a scalable, interpretable, and operational tool for real-time evaluation of re-identification risk.
2026
Data classification; Density-based analysis; Interpretability; Pattern generation; Re-identification risk
01 Pubblicazione su rivista::01a Articolo in rivista
A classification-based approach to the evaluation of personal identification risk for supposedly anonymous data / Bruni, R., Bianchi, G., Lorusso, P.. - In: INFORMATION PROCESSING & MANAGEMENT. - ISSN 1873-5371. - 64:1(2026). [10.1016/j.ipm.2026.105073]
File allegati a questo prodotto
Non ci sono file associati a questo prodotto.

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11573/1774796
 Attenzione

Attenzione! I dati visualizzati non sono stati sottoposti a validazione da parte dell'ateneo

Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus ND
  • ???jsp.display-item.citation.isi??? 0
social impact