TY - GEN
T1 - Feature Subset Selection for Learning Huge Configuration Spaces
T2 - 26th ACM International Systems and Software Product Line Conference, ASPLC 2022
AU - Acher, Mathieu
AU - Martin, Hugo
AU - Lesoil, Luc
AU - Blouin, Arnaud
AU - Jézéquel, Jean Marc
AU - Khelladi, Djamel Eddine
AU - Barais, Olivier
AU - Pereira, Juliana Alves
N1 - Publisher Copyright:
© 2022 ACM.
PY - 2022/9/12
Y1 - 2022/9/12
N2 - Linux kernels are used in a wide variety of appliances, many of them having strong requirements on the kernel size due to constraints such as limited memory or instant boot. With more than nine thousands of configuration options to choose from, developers and users of Linux actually spend significant effort to document, understand, and eventually tune (combinations of) options for meeting a kernel size. In this paper, we describe a large-scale endeavour automating this task and predicting a given Linux kernel binary size out of unmeasured configurations. We first experiment that state-of-the-art solutions specifically made for configurable systems such as performance-influence models cannot cope with that number of options, suggesting that software product line techniques may need to be adapted to such huge configuration spaces. We then show that tree-based feature selection can learn a model achieving low prediction errors over a reduced set of options. The resulting model, trained on 95 854 kernel configurations, is fast to compute, simple to interpret and even outperforms the accuracy of learning without feature selection.
AB - Linux kernels are used in a wide variety of appliances, many of them having strong requirements on the kernel size due to constraints such as limited memory or instant boot. With more than nine thousands of configuration options to choose from, developers and users of Linux actually spend significant effort to document, understand, and eventually tune (combinations of) options for meeting a kernel size. In this paper, we describe a large-scale endeavour automating this task and predicting a given Linux kernel binary size out of unmeasured configurations. We first experiment that state-of-the-art solutions specifically made for configurable systems such as performance-influence models cannot cope with that number of options, suggesting that software product line techniques may need to be adapted to such huge configuration spaces. We then show that tree-based feature selection can learn a model achieving low prediction errors over a reduced set of options. The resulting model, trained on 95 854 kernel configurations, is fast to compute, simple to interpret and even outperforms the accuracy of learning without feature selection.
U2 - 10.1145/3546932.3546997
DO - 10.1145/3546932.3546997
M3 - Conference contribution
AN - SCOPUS:85139157874
T3 - 26th ACM International Systems and Software Product Line Conference, SPLC 2022 - Proceedings
SP - 85
EP - 96
BT - 26th ACM International Systems and Software Product Line Conference, SPLC 2022 - Proceedings
A2 - Felfernig, Alexander
A2 - Fuentes, Lidia
A2 - Cleland-Huang, Jane
A2 - Assuncao, Wesley K.G.
A2 - Assuncao, Wesley K.G.
A2 - Falkner, Andreas
A2 - Azanza, Maider
A2 - Luaces, Miguel A. Rodriguez
A2 - Bhushan, Megha
A2 - Semini, Laura
A2 - Devroey, Xavier
A2 - Werner, Claudia Maria Lima
A2 - Seidl, Christoph
A2 - Le, Viet-Man
A2 - Horcas, Jose Miguel
PB - Association for Computing Machinery, Inc
Y2 - 12 September 2022 through 16 September 2022
ER -