High-dimensional dense embeddings have become central to modern Information Retrieval, but many dimensions are noisy or redundant. Recently proposed DIME (Dimension IMportance Estimation), provides query-dependent scores to identify informative components of embeddings. DIME relies on a costly grid search to select a priori a dimensionality for all the query corpus’s embeddings. Our work provides a statistically grounded criterion that directly identifies the optimal set of dimensions for each query at inference time. Experiments confirm achieving parity of effectiveness and reduces embedding size by an average of ∼50% across different models and datasets at inference time
Statistical Foundations of DIME: Risk Estimation for Practical Index Selection / D'Erasmo, G., Campagnano, C., Mallia, A., Brutti, P., Tonellotto, N., Silvestri, F.. - (2026), pp. 722-730. (19th Conference of the European Chapter of the Association for Computational Linguistics (EACL) Rabat; Morocco ) [10.18653/v1/2026.eacl-short.51].
Statistical Foundations of DIME: Risk Estimation for Practical Index Selection
D'Erasmo Giulio
;Campagnano Cesare;Brutti Pierpaolo;Tonellotto Nicola;Silvestri Fabrizio
2026
Abstract
High-dimensional dense embeddings have become central to modern Information Retrieval, but many dimensions are noisy or redundant. Recently proposed DIME (Dimension IMportance Estimation), provides query-dependent scores to identify informative components of embeddings. DIME relies on a costly grid search to select a priori a dimensionality for all the query corpus’s embeddings. Our work provides a statistically grounded criterion that directly identifies the optimal set of dimensions for each query at inference time. Experiments confirm achieving parity of effectiveness and reduces embedding size by an average of ∼50% across different models and datasets at inference timeI documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.


