An administrative archive holds the cases a health system has noticed, not the cases that exist, and the share it misses is rarely uniform: it is greatest in developing countries and, within a country or region, in its more deprived areas. Prevalence estimated from such counts therefore tends to be understated where disease is most common, with practical consequences, since these estimates govern intervention policy and the allocation of health resources. Existing corrections require either validation data, which are rarely available, or strong prior information about how completely each area reports, typically supplied by partitioning the areas into data quality groups fixed in advance. Both requirements are restrictive, and the second carries a particular risk, since an inaccurate partition propagates into the corrected estimates without any indication within the model that it has done so. We propose a nonparametric compound Poisson model in which a latent clustering structure for the reporting probabilities is estimated jointly with the remaining parameters, informed by an auxiliary variable associated with reporting. The clustering is thus an object of inference rather than an input: the number of data quality groups need not be specified in advance, and the uncertainty about the partition is carried through to the corrected estimates. Identifiability, not granted in models of this kind, is secured by an order constraint on the reporting probabilities together with prior information about the best reporting areas alone. The model is applied to Chronic Kidney Disease (CKD) prevalence in Apulia, Italy, one of the most deprived and heterogeneous Italian regions, where counts are expected to be considerably underreported, using unpublished administrative data covering 2011 to 2013. The estimates agree with previous findings and with expert expectation, and display coherent spatial patterns. Against existing approaches, on simulated data and on early neonatal mortality data from Minas Gerais, Brazil, the model proves accurate and especially suitable when prior information about data quality is itself imperfect. A final chapter considers the use of such a map. Taking the DANTE case-finding intervention, which raised detected CKD prevalence in Apulian primary care by 73% in six months, a five-year costing puts the dialysis that earlier detection would avert at some €1,500 per patient-year among diabetic patients. It also finds that population size outweighs the underreporting signal in determining where case-finding would be most productive, so that the estimates are better suited to identifying where recording requires attention than to reallocating clinical effort.

A Bayesian nonparametric approach to correct for underreporting in count data / Pasculli, G.. - (2022 Oct 22).

A Bayesian nonparametric approach to correct for underreporting in count data

PASCULLI, GIUSEPPE
22/10/2022

Abstract

An administrative archive holds the cases a health system has noticed, not the cases that exist, and the share it misses is rarely uniform: it is greatest in developing countries and, within a country or region, in its more deprived areas. Prevalence estimated from such counts therefore tends to be understated where disease is most common, with practical consequences, since these estimates govern intervention policy and the allocation of health resources. Existing corrections require either validation data, which are rarely available, or strong prior information about how completely each area reports, typically supplied by partitioning the areas into data quality groups fixed in advance. Both requirements are restrictive, and the second carries a particular risk, since an inaccurate partition propagates into the corrected estimates without any indication within the model that it has done so. We propose a nonparametric compound Poisson model in which a latent clustering structure for the reporting probabilities is estimated jointly with the remaining parameters, informed by an auxiliary variable associated with reporting. The clustering is thus an object of inference rather than an input: the number of data quality groups need not be specified in advance, and the uncertainty about the partition is carried through to the corrected estimates. Identifiability, not granted in models of this kind, is secured by an order constraint on the reporting probabilities together with prior information about the best reporting areas alone. The model is applied to Chronic Kidney Disease (CKD) prevalence in Apulia, Italy, one of the most deprived and heterogeneous Italian regions, where counts are expected to be considerably underreported, using unpublished administrative data covering 2011 to 2013. The estimates agree with previous findings and with expert expectation, and display coherent spatial patterns. Against existing approaches, on simulated data and on early neonatal mortality data from Minas Gerais, Brazil, the model proves accurate and especially suitable when prior information about data quality is itself imperfect. A final chapter considers the use of such a map. Taking the DANTE case-finding intervention, which raised detected CKD prevalence in Apulian primary care by 73% in six months, a five-year costing puts the dialysis that earlier detection would avert at some €1,500 per patient-year among diabetic patients. It also finds that population size outweighs the underreporting signal in determining where case-finding would be most productive, so that the estimates are better suited to identifying where recording requires attention than to reallocating clinical effort.
22-ott-2022
Pesce, Francesco
File allegati a questo prodotto
File Dimensione Formato  
Tesi_dottorato_Pasculli.pdf

accesso aperto

Note: A-Bayesian-Nonparametric-Approach-to-Correct-for-Underreporting-in-Count-Data
Tipologia: Tesi di dottorato
Licenza: Creative commons
Dimensione 11.79 MB
Formato Adobe PDF
11.79 MB Adobe PDF

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11573/1773210
Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus ND
  • ???jsp.display-item.citation.isi??? ND
social impact