TY - GEN
T1 - Random forests of very fast decision trees on GPU for mining evolving big data streams
AU - Marron, Diego
AU - Bifet, Albert
AU - De Francisci Morales, Gianmarco
N1 - Publisher Copyright:
© 2014 The Authors and IOS Press.
PY - 2014/1/1
Y1 - 2014/1/1
N2 - Random Forest is a classical ensemble method used to improve the performance of single tree classifiers. It is able to obtain superior performance by increasing the diversity of the single classifiers. However, in the more challenging context of evolving data streams, the classifier has also to be adaptive and work under very strict constraints of space and time. Furthermore, the computational load of using a large number of classifiers can make its application extremely expensive. In this work, we present a method for building Random Forests that use Very Fast Decision Trees for data streams on GPUs. We show how this method can benefit from the massive parallel architecture of GPUs, which are becoming an efficient hardware alternative to large clusters of computers. Moreover, our algorithm minimizes the communication between CPU and GPU by building the trees directly inside the GPU. We run an empirical evaluation and compare our method to two well know machine learning frameworks, VFML and MOA. Random Forests on the GPU are at least 300x faster while maintaining a similar accuracy.
AB - Random Forest is a classical ensemble method used to improve the performance of single tree classifiers. It is able to obtain superior performance by increasing the diversity of the single classifiers. However, in the more challenging context of evolving data streams, the classifier has also to be adaptive and work under very strict constraints of space and time. Furthermore, the computational load of using a large number of classifiers can make its application extremely expensive. In this work, we present a method for building Random Forests that use Very Fast Decision Trees for data streams on GPUs. We show how this method can benefit from the massive parallel architecture of GPUs, which are becoming an efficient hardware alternative to large clusters of computers. Moreover, our algorithm minimizes the communication between CPU and GPU by building the trees directly inside the GPU. We run an empirical evaluation and compare our method to two well know machine learning frameworks, VFML and MOA. Random Forests on the GPU are at least 300x faster while maintaining a similar accuracy.
U2 - 10.3233/978-1-61499-419-0-615
DO - 10.3233/978-1-61499-419-0-615
M3 - Conference contribution
AN - SCOPUS:84923205468
T3 - Frontiers in Artificial Intelligence and Applications
SP - 615
EP - 620
BT - ECAI 2014 - 21st European Conference on Artificial Intelligence, Including Prestigious Applications of Intelligent Systems, PAIS 2014, Proceedings
A2 - Schaub, Torsten
A2 - Friedrich, Gerhard
A2 - O'Sullivan, Barry
PB - IOS Press BV
T2 - 21st European Conference on Artificial Intelligence, ECAI 2014
Y2 - 18 August 2014 through 22 August 2014
ER -