login

Distribution Problems in Clustering

Elsevier eBooksPublished 1 January 1977
J. A. Hartigan
Citations101

TL;DR

A statistical problem that is encountered in deciding which of the many clusters presented by algorithms are real is discussed, which requires the asymptotic theory to be validated by Monte Carlo experiments.

Abstract

The very large growth in clustering techniques and applications is not yet supported by development of statistical theory by which the clustering results are evaluated. A number of branches of statistics are relevant to clustering, namely, discriminant analysis, eigenvector analysis, analysis of variance, multiple comparisons, density estimation, contingency tables, piecewise fitting, and regression. These are all areas where the techniques are used in evaluating clusters or where clustering operations occur. This chapter discusses a statistical problem that is encountered in deciding which of the many clusters presented by algorithms are real. There is no easy generally applicable definition of real. A data cluster is real if it corresponds to one of the population clusters. The mixture techniques, k-means, single linkage, complete linkage and other common algorithms are examined to give measures of the reality of their clusters. A reasonable significance testing procedure requires the asymptotic theory to be validated by Monte Carlo experiments.

Keywords

Computer ScienceMathematics