TY - GEN
T1 - Predictive models of hard drive failures based on operational data
AU - Aussel, Nicolas
AU - Jaulin, Samuel
AU - Gandon, Guillaume
AU - Petetin, Yohan
AU - Fazli, Eriza
AU - Chabridon, Sophie
N1 - Publisher Copyright:
© 2017 IEEE.
PY - 2017/1/1
Y1 - 2017/1/1
N2 - Hard drives are an essential component of modern data storage. In order to reduce the risk of data loss, hard drive failure prediction methods using the Self-Monitoring, Analysis and Reporting Technology attributes have been proposed. However, these methods were developed from datasets not necessarily representative of operational systems. In this paper, we consider the Backblaze public dataset, a recent operational dataset from over 47,000 drives, exhibiting hard drive heterogeneity with 81 models from 5 manufacturers, an extremely unbalanced ratio of 5000:1 between healthy and failure samples and a realworld loosely controlled environment. We observe that existing predictive models no longer perform sufficiently well on this dataset. We therefore selected machine learning classification methods able to deal with a very unbalanced training set, namely SVM, RF and GBT, and adapted them to the specific constraints of hard drive failure prediction. Our results reach over 95% precision and 67% recall on a one year real-world public dataset of over 12 million records with only 2586 failures.
AB - Hard drives are an essential component of modern data storage. In order to reduce the risk of data loss, hard drive failure prediction methods using the Self-Monitoring, Analysis and Reporting Technology attributes have been proposed. However, these methods were developed from datasets not necessarily representative of operational systems. In this paper, we consider the Backblaze public dataset, a recent operational dataset from over 47,000 drives, exhibiting hard drive heterogeneity with 81 models from 5 manufacturers, an extremely unbalanced ratio of 5000:1 between healthy and failure samples and a realworld loosely controlled environment. We observe that existing predictive models no longer perform sufficiently well on this dataset. We therefore selected machine learning classification methods able to deal with a very unbalanced training set, namely SVM, RF and GBT, and adapted them to the specific constraints of hard drive failure prediction. Our results reach over 95% precision and 67% recall on a one year real-world public dataset of over 12 million records with only 2586 failures.
U2 - 10.1109/ICMLA.2017.00-92
DO - 10.1109/ICMLA.2017.00-92
M3 - Conference contribution
AN - SCOPUS:85048477933
T3 - Proceedings - 16th IEEE International Conference on Machine Learning and Applications, ICMLA 2017
SP - 619
EP - 625
BT - Proceedings - 16th IEEE International Conference on Machine Learning and Applications, ICMLA 2017
A2 - Chen, Xuewen
A2 - Luo, Bo
A2 - Luo, Feng
A2 - Palade, Vasile
A2 - Wani, M. Arif
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 16th IEEE International Conference on Machine Learning and Applications, ICMLA 2017
Y2 - 18 December 2017 through 21 December 2017
ER -