Passer à la navigation principale Passer à la recherche Passer au contenu principal

Feature Scoring using Tree-Based Ensembles for Evolving Data Streams

  • Heitor Murilo Gomes
  • , Rodrigo Fernandes De Mello
  • , Bernhard Pfahringer
  • , Albert Bifet
  • University of Waikato
  • Department of Computer Science (ICMC)
  • University of São Paulo

Résultats de recherche: Le chapitre dans un livre, un rapport, une anthologie ou une collectionContribution à une conférenceRevue par des pairs

13 Citations (Scopus)

Résumé

Assigning scores to individual features is a popular method for estimating the relevance of features in supervised learning. An accurate feature score estimation provides essential insights in sensitive domains, which is decisive to explain how features influence a given decision, contributing to the interpretability of the model. Learning from streaming data adds several challenges to machine learning tasks, including limited resources and changes to the underlying data distribution (i.e., evolving data streams). In this work, we introduce and analyze methods to efficiently estimate the Mean Decrease in Impurity (MDI) and COVER measures using ensembles of incremental decision trees. To achieve current scores in evolving data streams, we employ tree-ensembles that incorporate active drift detection. Experimental results show how MDI and COVER can be used to track the feature scores when their importance to the ensemble model shift over time. On top of that, we present the impact on the feature scores when the learning problem includes a non-negligible verification latency for the arrival of the labels. We also present a counter-intuitive experiment using a standard benchmark dataset where the feature scores correctly illustrate the importance of two features to the ensemble model. However, these features are prioritized due to biased split decisions, and in their absence, the model increases in predictive performance. We conclude that the presented measures can be used to understand the impact of features in the ensemble model better, still, such measures should be used with caution as they are limited by the underlying tree building and ensemble model biases.

langue originaleAnglais
titreProceedings - 2019 IEEE International Conference on Big Data, Big Data 2019
rédacteurs en chefChaitanya Baru, Jun Huan, Latifur Khan, Xiaohua Tony Hu, Ronay Ak, Yuanyuan Tian, Roger Barga, Carlo Zaniolo, Kisung Lee, Yanfang Fanny Ye
EditeurInstitute of Electrical and Electronics Engineers Inc.
Pages761-769
Nombre de pages9
ISBN (Electronique)9781728108582
Les DOIs
étatPublié - 1 déc. 2019
Modification externeOui
Evénement2019 IEEE International Conference on Big Data, Big Data 2019 - Los Angeles, États-Unis
Durée: 9 déc. 201912 déc. 2019

Série de publications

NomProceedings - 2019 IEEE International Conference on Big Data, Big Data 2019

Une conférence

Une conférence2019 IEEE International Conference on Big Data, Big Data 2019
Pays/TerritoireÉtats-Unis
La villeLos Angeles
période9/12/1912/12/19

Empreinte digitale

Examiner les sujets de recherche de « Feature Scoring using Tree-Based Ensembles for Evolving Data Streams ». Ensemble, ils forment une empreinte digitale unique.

Contient cette citation