We introduce a load-aware selective embedding-pruning method to be used in dense retrieval systems. Our method adapts the embedding dimensionality at inference time based on per-query deadlines and the current queue occupancy. Under light load, full-dimensional embeddings maximise effectiveness; as load increases, dimensionality is adaptively reduced to maintain throughput and avoid query drops. Our approach is orthogonal to the specific dimensionality reduction strategy employed and operates independently of the embedding model. We evaluate our approach using both PCA-reduced embeddings and nested Matryoshka embeddings. Empirical results show that our load-aware strategy consistently achieves a better effectiveness-efficiency trade-off than static dimensionality reduction baselines across varying load conditions. As a by-product of this study, our experiments show that under selective pruning, the PCA-based approach consistently outperforms Matryoshka, indicating that specialised multi-representation training is not strictly required for robust load balancing in dense retrieval.
Load-sensitive Selective Pruning in Dense Retrieval / Calugaru, M.D., Siciliano, F., Pezzuti, F., Tonellotto, N., Silvestri, F.. - (2026), pp. 3630-3634. (49th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR 2026) Melbourne, Australia ) [10.1145/3805712.3809875].
Load-sensitive Selective Pruning in Dense Retrieval
Maria Diana Calugaru;Federico Siciliano;Nicola Tonellotto;Fabrizio Silvestri
2026
Abstract
We introduce a load-aware selective embedding-pruning method to be used in dense retrieval systems. Our method adapts the embedding dimensionality at inference time based on per-query deadlines and the current queue occupancy. Under light load, full-dimensional embeddings maximise effectiveness; as load increases, dimensionality is adaptively reduced to maintain throughput and avoid query drops. Our approach is orthogonal to the specific dimensionality reduction strategy employed and operates independently of the embedding model. We evaluate our approach using both PCA-reduced embeddings and nested Matryoshka embeddings. Empirical results show that our load-aware strategy consistently achieves a better effectiveness-efficiency trade-off than static dimensionality reduction baselines across varying load conditions. As a by-product of this study, our experiments show that under selective pruning, the PCA-based approach consistently outperforms Matryoshka, indicating that specialised multi-representation training is not strictly required for robust load balancing in dense retrieval.I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.


