login

K-means properties on six clustering benchmark datasets

Applied IntelligencePublished 26 July 2018
Pasi Fränti, Sami Sieranoja
Citations511
SJR quartileQ2
SJR score0.93
SNIP1.21

TL;DR

The results show that overlap is critical, and that k-means starts to work effectively when the overlap reaches 4% level.

Abstract

This paper has two contributions. First, we introduce a clustering basic benchmark. Second, we study the performance of k-means using this benchmark. Specifically, we measure how the performance depends on four factors: (1) overlap of clusters, (2) number of clusters, (3) dimensionality, and (4) unbalance of cluster sizes. The results show that overlap is critical, and that k-means starts to work effectively when the overlap reaches 4% level.

Keywords

Computer Science