TY - GEN
T1 - A data set oriented approach for clustering algorithm selection
AU - Halkich, Maria
AU - Vazirgiannis, Michalis
N1 - Publisher Copyright:
© Springer-Verlag Berlin Heidelberg 2001.
PY - 2001/1/1
Y1 - 2001/1/1
N2 - In the last years the availability of huge transactional and experimental data sets and the arising requirements for data mining created needs for clustering algorithms that scale and can be applied in diverse domains. Thus, a variety of algorithms have been proposed which have application in different fields and may result in different partitioning of a data set, depending on the specific clustering criterion used. Moreover, since clustering is an unsupervised process, most of the algorithms are based on assumptions in order to define a partitioning of a data set. It is then obvious that in most applications the final clustering scheme requires some sort of evaluation. In this paper we present a clustering validity procedure, which taking in account the inherent features of a data set evaluates the results of different clustering algorithms applied to it. A validity index, S_Dbw, is defined according to wellknown clustering criteria so as to enable the selection of the algorithm providing the best partitioning of a data set. We evaluate the reliability of our approach both theoretically and experimentally, considering three representative clustering algorithms ran on synthetic and real data sets. It performed favorably in all studies, giving an indication of the algorithm that is suitable for the considered application.
AB - In the last years the availability of huge transactional and experimental data sets and the arising requirements for data mining created needs for clustering algorithms that scale and can be applied in diverse domains. Thus, a variety of algorithms have been proposed which have application in different fields and may result in different partitioning of a data set, depending on the specific clustering criterion used. Moreover, since clustering is an unsupervised process, most of the algorithms are based on assumptions in order to define a partitioning of a data set. It is then obvious that in most applications the final clustering scheme requires some sort of evaluation. In this paper we present a clustering validity procedure, which taking in account the inherent features of a data set evaluates the results of different clustering algorithms applied to it. A validity index, S_Dbw, is defined according to wellknown clustering criteria so as to enable the selection of the algorithm providing the best partitioning of a data set. We evaluate the reliability of our approach both theoretically and experimentally, considering three representative clustering algorithms ran on synthetic and real data sets. It performed favorably in all studies, giving an indication of the algorithm that is suitable for the considered application.
U2 - 10.1007/3-540-44794-6_14
DO - 10.1007/3-540-44794-6_14
M3 - Conference contribution
AN - SCOPUS:47649129920
SN - 9783540425342
T3 - Lecture Notes in Computer Science
SP - 165
EP - 179
BT - Principles of Data Mining and Knowledge Discovery - 5th European Conference, PKDD 2001, Proceedings
A2 - De Raedt, Luc
A2 - Siebes, Arno
PB - Springer Verlag
T2 - 5th European Conference on Principles of Data Mining and Knowledge Discovery, PKDD 2001
Y2 - 3 September 2001 through 5 September 2001
ER -