article · International Journal of Advanced Statistics and Probability
This study improves one of the initialization methods for the k-means clustering algorithm based on a rough set neighbourhood model to enhance performance in noisy datasets. The method involves data normalization, obtaining a neighbourhood threshold based on the 0.25th trimmed mean of pairwise Minkowski distances, calculating cohesion and coupling degrees of the neighbourhoods and between them re-spectively, and obtaining the initial cluster centres as the k points having maximum cohesion degrees with minimum coupling degrees among themselves. The approach was evaluated on six datasets using Silhouette, Davies–Bouldin, Calinski–Harabasz, and Dunn–Hubert indices in comparison with an existing method. Results showed that the improved method outperformed the existing method on noisy datasets, achieving higher Silhouette and Dunn–Hubert scores, and lower Davies–Bouldin values, with a slight reduction in Calinski–Harabasz index in one of the datasets. On the non-noisy datasets, the two methods were at par in all four performance indices. With the improved performance, showing that the improved method enhanced the stability and robustness of k-means clustering in the presence of noisy data, it can be recommended for clustering noisy datasets such as gene expression, image, and signal datasets.
This page summarises published work. The authoritative version sits with the publisher.
DOI: 10.14419/69dvcw11
Is something wrong with this record? Report it or request removal.
Discussion
Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.
No discussion yet. Open the first thread.
New to MARATTO™? Create a free account.