Compressive genomics leverages compressed data representations to enhance the efficiency of bioinformatics tasks like sequence comparison and search. Surprisingly, the fundamental operation of pattern matching on large DNA sequence collections remains unexplored in the realm of genomic analysis. However, distributed systems like Spark offer the scalability necessary to process increasingly large genomic datasets efficiently. We present the first Spark-based implementation of the FM-Index and Compressed Boyer-Moore (CBM) algorithms, evaluating their performance and providing insights into their advantages for large-scale bioinformatics applications. A comprehensive experimental study demonstrates clear performance gains over uncompressed approaches. Furthermore, we introduce SparkGeco, a distributed compressive genomics software library designed to simplify the integration of FM-Index and CBM algorithms into DNA sequence analysis pipelines within Apache Spark, thus supporting the development of efficient and scalable genomic analysis workflows. This work provides a concrete step towards high-performance, data-centric eScience solutions in computational biology.

Distributed compressive genomics: Fundamental pattern matching primitives via spark / Rocco, L.D., Ferraro Petrillo, U., Giancarlo, R., Cattaneo, G.. - In: FUTURE GENERATION COMPUTER SYSTEMS. - ISSN 0167-739X. - 176:(2026). [10.1016/j.future.2025.108169]

Distributed compressive genomics: Fundamental pattern matching primitives via spark

Rocco, Lorenzo Di;Ferraro Petrillo, Umberto;Giancarlo, Raffaele;Cattaneo, Giuseppe
2026

Abstract

Compressive genomics leverages compressed data representations to enhance the efficiency of bioinformatics tasks like sequence comparison and search. Surprisingly, the fundamental operation of pattern matching on large DNA sequence collections remains unexplored in the realm of genomic analysis. However, distributed systems like Spark offer the scalability necessary to process increasingly large genomic datasets efficiently. We present the first Spark-based implementation of the FM-Index and Compressed Boyer-Moore (CBM) algorithms, evaluating their performance and providing insights into their advantages for large-scale bioinformatics applications. A comprehensive experimental study demonstrates clear performance gains over uncompressed approaches. Furthermore, we introduce SparkGeco, a distributed compressive genomics software library designed to simplify the integration of FM-Index and CBM algorithms into DNA sequence analysis pipelines within Apache Spark, thus supporting the development of efficient and scalable genomic analysis workflows. This work provides a concrete step towards high-performance, data-centric eScience solutions in computational biology.
2026
Compressive genomics; Data compression; Distributed systems; FM-Index; Pattern matching; Spark
01 Pubblicazione su rivista::01a Articolo in rivista
Distributed compressive genomics: Fundamental pattern matching primitives via spark / Rocco, L.D., Ferraro Petrillo, U., Giancarlo, R., Cattaneo, G.. - In: FUTURE GENERATION COMPUTER SYSTEMS. - ISSN 0167-739X. - 176:(2026). [10.1016/j.future.2025.108169]
File allegati a questo prodotto
Non ci sono file associati a questo prodotto.

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11573/1774939
 Attenzione

Attenzione! I dati visualizzati non sono stati sottoposti a validazione da parte dell'ateneo

Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus 1
  • ???jsp.display-item.citation.isi??? 1
social impact