TY - JOUR
T1 - Spark
T2 - sparse hierarchical energy minimization for scalable prediction of RNA pseudoknots
AU - Gray, Mateo
AU - Will, Sebastian
AU - Jabbari, Hosna
N1 - Publisher Copyright:
© The Author(s) 2026. Published by Oxford University Press. This is an Open Access article distributed under the terms of the Creative Commons Attribution License (https://creativecommons.org/licenses/by/4.0/), which permits unrestricted reuse, distribution, and reproduction in any medium, provided the original work is properly cited.
PY - 2026/5/1
Y1 - 2026/5/1
N2 - Motivation: The biological functions of RNAs are tightly connected to their specific RNA structures. As experimental techniques to determine high-accuracy structures are costly and time-consuming, computational prediction approaches became indispensable for biological RNA research; most notably, the prediction of minimum free energy secondary structures. Pseudoknots are prevalent, highly significant structural motifs, yet they are commonly ignored to achieve acceptable efficiency. Existing reliable pseudoknot prediction methods typically have prohibitive complexity. A route to fast scalable pseudoknot prediction was suggested with HFold following the hierarchical folding hypothesis. Recent successful sparsification of the CCJ pseudoknot prediction algorithm in Knotty promises a further boost by introducing this technique to hierarchical folding. Results: We introduce Spark, a sparsified algorithm for predicting pseudoknotted RNA structures. Spark predicts exactly the same minimum-energy structures as its predecessor HFold in the accurate HotKnots 2.0 energy model for pseudoknots. While sparsification maintains exact energy minimization and theoretical complexity, it strongly improves the time and space consumption over HFold. We benchmarked the performance of Spark against HFold and, as a pseudoknot-free baseline, RNAfold. Compared with HFold, Spark substantially reduces both run time and memory usage, while achieving run times close to RNAfold. Across all tested sequence lengths, Spark used the least memory and consistently ran faster than HFold. Conclusion: Combining sparsification and hierarchical folding in Spark results in an remarkably fast and memory-efficient tool for the accurate prediction of pseudoknotted RNA structures. Consequently, Spark practically enables pseudoknot prediction in large scale and even for very long RNA sequences. Availability: Spark software is available on Github (https://github.com/TheCOBRALab/Spark), with a permanent archive of the software and results deposited on Zenodo (https://doi.org/10.5281/zenodo.19073315).
AB - Motivation: The biological functions of RNAs are tightly connected to their specific RNA structures. As experimental techniques to determine high-accuracy structures are costly and time-consuming, computational prediction approaches became indispensable for biological RNA research; most notably, the prediction of minimum free energy secondary structures. Pseudoknots are prevalent, highly significant structural motifs, yet they are commonly ignored to achieve acceptable efficiency. Existing reliable pseudoknot prediction methods typically have prohibitive complexity. A route to fast scalable pseudoknot prediction was suggested with HFold following the hierarchical folding hypothesis. Recent successful sparsification of the CCJ pseudoknot prediction algorithm in Knotty promises a further boost by introducing this technique to hierarchical folding. Results: We introduce Spark, a sparsified algorithm for predicting pseudoknotted RNA structures. Spark predicts exactly the same minimum-energy structures as its predecessor HFold in the accurate HotKnots 2.0 energy model for pseudoknots. While sparsification maintains exact energy minimization and theoretical complexity, it strongly improves the time and space consumption over HFold. We benchmarked the performance of Spark against HFold and, as a pseudoknot-free baseline, RNAfold. Compared with HFold, Spark substantially reduces both run time and memory usage, while achieving run times close to RNAfold. Across all tested sequence lengths, Spark used the least memory and consistently ran faster than HFold. Conclusion: Combining sparsification and hierarchical folding in Spark results in an remarkably fast and memory-efficient tool for the accurate prediction of pseudoknotted RNA structures. Consequently, Spark practically enables pseudoknot prediction in large scale and even for very long RNA sequences. Availability: Spark software is available on Github (https://github.com/TheCOBRALab/Spark), with a permanent archive of the software and results deposited on Zenodo (https://doi.org/10.5281/zenodo.19073315).
UR - https://www.scopus.com/pages/publications/105038670755
U2 - 10.1093/bioinformatics/btag194
DO - 10.1093/bioinformatics/btag194
M3 - Article
C2 - 42015361
AN - SCOPUS:105038670755
SN - 1367-4803
VL - 42
JO - Bioinformatics
JF - Bioinformatics
IS - 5
M1 - btag194
ER -