The Practice of Cluster Analysis
Journal of ClassificationPublished 1 June 2006
J. R. Kettenring
Citations195
SJR quartileQ1
SJR score0.71
SNIP1.33
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
The goal of this article is to document this growth, characterize current usage, illustrate the breadth of applications via examples, highlight both good and risky practices, and suggest some research priorities.
Abstract
Cluster analysis is one of the main methodologies for analyzing multivariate data. Its use is widespread and growing rapidly. The goal of this article is to document this growth, characterize current usage, illustrate the breadth of applications via examples, highlight both good and risky practices, and suggest some research priorities.
Keywords
Computer ScienceMathematics
The Elements of Statistical Learning: Data Mining, Inference, and Prediction
19,345 Citations2013Trevor Hastie, Robert Tibshirani +1 more
Proceedings of the National Academy of SciencesCluster analysis and display of genome-wide expression patterns
16,395 Citations1998Michael B. Eisen, Paul T. Spellman +2 more
A system of cluster analysis for genome-wide expression data from DNA microarray hybridization is described that uses standard statistical algorithms to arrange genes according to similarity in pattern of gene expression, finding in the budding yeast Saccharomyces cerevisiae that clustering gene expression data groups together efficiently genes of known similar function.
Encyclopedia of Statistics in Behavioral SciencePrincipal Component Analysis
14,549 Citations2005Ian T. Jolliffe
BiometricsApplied Multivariate Statistical Analysis.
11,424 Citations1988Andrea Johnson, Dean W. Wichern
Choice Reviews OnlineFinding groups in data: an introduction to cluster analysis
10,614 Citations1991
Wiley series in probability and statisticsLinear Statistical Inference and its Applications
10,509 Citations1973C. Radhakrishna Rao
Journal of Business and Economic StatisticsAn Introduction to Multivariate Statistical Analysis
9,232 Citations1986Robb J. Muirhead, T. W. Anderson
Cluster Analysis
9,195 Citations1974Barry J. Everitt, Sabine Landau +1 more
GeneticsInference of Population Structure Using Multilocus Genotype Data: Linked Loci and Correlated Allele Frequencies
8,053 Citations2003Daniel Falush, Matthew Stephens +1 more
Extensions to the method of Pritchard et al. for inferring population structure from multilocus genotype data are described and methods that allow for linkage between loci are developed, which allows identification of subtle population subdivisions that were not detectable using the existing method.
IEEE Transactions on Neural NetworksSurvey of Clustering Algorithms
6,154 Citations2005Rui Xu, D. WunschII
Clustering algorithms for data sets appearing in statistics, computer science, and machine learning are surveyed, and their applications in some benchmark data sets, the traveling salesman problem, and bioinformatics, a new field attracting intensive efforts are illustrated.
Journal of the American Statistical AssociationModel-Based Clustering, Discriminant Analysis, and Density Estimation
4,282 Citations2002Chris Fraley, Adrian E. Raftery
This work reviews a general methodology for model-based clustering that provides a principled statistical approach to important practical questions that arise in cluster analysis, such as how many clusters are there, which clustering method should be used, and how should outliers be handled.
Minds at UW (University of Wisconsin)Semi-Supervised Learning Literature Survey
3,868 Citations2005Xiaojin Zhu
The study clearly indicates that the common practice of stripwise precommercial thinning is unjustified, and the justification of heavy 'chessboard' thinning (with pruning) depends on whether the potential reduction in rotation length and the improvement in wood quality outweigh the discounted costs of pre-commercial thinning and selection and pruning of crop trees.
BiometricsModel-Based Gaussian and Non-Gaussian Clustering
2,350 Citations1993Jeffrey D. Banfield, Adrian E. Raftery
The classification maximum likelihood approach is sufficiently general to encompass many current clustering algorithms, including those based on the sum of squares criterion and on the criterion of Friedman and Rubin (1967), but it is restricted to Gaussian distributions and it does not allow for noise.
BioinformaticsPrincipal component analysis for clustering gene expression data
1,294 Citations2001Ka Yee Yeung, Walter L. Ruzzo
The empirical study showed that clustering with the PCs instead of the original variables does not necessarily improve, and often degrades, cluster quality, and would not recommend PCA before clustering except in special circumstances.
IEEE Transactions on Knowledge and Data EngineeringCluster analysis for gene expression data: a survey
1,252 Citations2004Daxin Jiang, Chun Tang +1 more
This paper divides cluster analysis for gene expression data into three categories, presents specific challenges pertinent to each clustering category and introduces several representative approaches, and suggests the promising trends in this field.
Drug SafetyPrinciples of Data Mining
1,034 Citations2007David J. Hand
The book consists of three sections and provides a tutorial overview of the principles underlying data mining algorithms and their application, and shows how all of the preceding analysis fits together when applied to real-world data mining problems.
Journal of the American Statistical AssociationMethods for Statistical Data Analysis of Multivariate Observations.
900 Citations1978Pranab Kumar Sen, R. Gnanadesikan
BioinformaticsModel-based clustering and data transformations for gene expression data
888 Citations2001Ka Yee Yeung, Chris Fraley +3 more
A probabilistic framework for semi-supervised clustering
819 Citations2004Sugato Basu, Mikhail Bilenko +1 more
A probabilistic model for semi-supervised clustering based on Hidden Markov Random Fields (HMRFs) that provides a principled framework for incorporating supervision into prototype-based clustering and experimental results demonstrate the advantages of the proposed framework.
ScienceGenetic Structure of the Purebred Domestic Dog
734 Citations2004Heidi G. Parker, Lisa V. Kim +8 more
This work identified four genetic clusters, which predominantly contained breeds with similar geographic origin, morphology, or role in human activities, which will aid studies of the genetics of phenotypic breed differences.
Journal of the American Statistical AssociationHyperdimensional Data Analysis Using Parallel Coordinates
661 Citations1990Edward J. Wegman
The basic algorithm for parallel coordinates is laid out and a discussion of its properties as a projective transformation is given, and several duality results are discussed along with their interpretations as data analysis tools.
Proceedings of the National Academy of SciencesGene expression patterns in human embryonic stem cells and human pluripotent germ cell tumors
648 Citations2003Jamie M. Sperger, Xin Chen +8 more
The gene expression patterns of human ES cell lines showed many similarities with the human embryonal carcinoma cell samples and more distantly with the seminoma samples, and 895 genes were identified that are candidates for involvement in the maintenance of a pluripotent, undifferentiated phenotype.
Journal of Cross-Cultural PsychologyToward a Geography of Personality Traits
631 Citations2004Jüri Allïk, Robert R. McCrae
Journal of the Royal Statistical Society Series B (Statistical Methodology)Clustering Objects on Subsets of Attributes (with Discussion)
427 Citations2004Jerome H. Friedman, Jacqueline J. Meulman
A new procedure is proposed for clustering attribute value data that encourages those algorithms to detect automatically subgroups of objects that preferentially cluster on subsets of the attribute variables rather than on all of them simultaneously.
Applied Multivariate Statistics in Geohydrology and Related Sciences
401 Citations1998Charles E. Brown
An introduction to general statistical and multivariate concepts and other approaches to explore multivariate data multivariate measures of space, time and distance multivariateData preparation, plotting and conclusions.
BiometrikaRobust estimation and outlier detection with correlation coefficients
386 Citations1975Susan J. Devlin, R. Gnanadesikan +1 more
Multivariate Statistics for the Environmental Sciences
348 Citations2003Peter J. Shaw
The overall aim of the book is to introduce inexperienced users gently to the multivariate analytical tools available to them.
Journal of ClassificationEnhanced Model-Based Clustering, Density Estimation, and Discriminant Analysis Software: MCLUST
334 Citations2003Chris Fraley, Adrian E. Raftery
MCLUST is a software package for model-based clustering, density estimation and discriminant analysis interfaced to the S-PLUS commercial software and the R language that implements parameterized Gaussian hierarchical clustering algorithms and the EM algorithm for parameterizedGaussian mixture models with the possible addition of a Poisson noise term.
Journal of Computational and Graphical StatisticsInteractive High-Dimensional Data Visualization
332 Citations1996Andreas Buja, Dianne Cook +1 more
A rudimentary taxonomy of interactive data visualization is proposed based on a triad of data analytic tasks: finding Gestalt, posing queries, and making comparisons; namely, high-dimensional projections, linked scatterplot brushing, and matrices of conditional plots.
Journal of the Royal Statistical Society Series C (Applied Statistics)On Using Principal Components Before Separating a Mixture of Two Multivariate Normal Distributions
308 Citations1983Wei-Chien Chang
SIAM Journal on Scientific ComputingAlgorithms for Model-Based Gaussian Hierarchical Clustering
268 Citations1998Chris Fraley
It is shown how the structure of the Gaussian model can be exploited to yield efficient algorithms for agglomerative hierarchical clustering.
Psychological MethodsLocal Optima in K-Means Clustering: What You Don't Know May Hurt You.
263 Citations2003Douglas Steinley
The results suggest the need for some strategy to study the local optima problem for a specific data set or to identify methods for finding "good" starting values that might lead to the best solutions possible.
BiometricsTight Clustering: A Resampling‐Based Approach for Identifying Stable and Tight Patterns in Data
225 Citations2005George C. Tseng, Wing Hung Wong
A method for clustering that produces tight and stable clusters without forcing all points into clusters is proposed and applied to analyze a set of expression profiles in the study of embryonic stem cells.
Journal of ClassificationEstimating the Cluster Tree of a Density by Analyzing the Minimal Spanning Tree of a Sample
182 Citations2003Werner Stuetzle
In this work,unt pruning, a new clustering method that attempts to find modes of a density by analyzing the minimal spanning tree of a sample by exploiting the connection between the minimal spans tree and nearest neighbor density is introduced.
Journal of ClassificationWeighting and selection of variables for cluster analysis
164 Citations1995R. Gnanadesikan, J. R. Kettenring +1 more
This paper reports on the performance of nine methods on eight “leading case” simulated and real sets of data and demonstrates shortcomings of weighting based on the standard deviation or range as well as other more complex schemes in the literature.
Statistics in Archaeology
134 Citations2010Michael Baxter
Statistical ScienceStatistical Challenges in Functional Genomics
123 Citations2003Paola Sebastiani, Emanuela Gussoni +2 more
The foundations of this technology, known as microarrays, are reviewed and the statistical challenges posed by the analysis of microarray data are described.
The GerontologistPatterns and Impact of Comorbidity and Multimorbidity Among Community-Resident American Indian Elders
122 Citations2003Robert John, Dave S. Kerby +1 more
PoeticsMedia repertoires of selective audiences: the impact of status, gender, and age on media use
111 Citations2003Kees van Rees, Koen van Eijck
Journal of Statistical PhysicsCluster Analysis of Gene Expression Data
83 Citations2003Eytan Domany
This review provides a very basic introduction to the subject of gene expression and clustering, aimed at a physics audience with no prior knowledge of either gene expression or clustering methods.
Human Brain MappingCluster analysis of fMRI data using dendrogram sharpening
57 Citations2003Larissa Stanberry, Rajesh Nandy +1 more
The dendrogram sharpening method, combined with a hierarchical clustering algorithm, is used in this work to identify modality regions, which are, in essence, areas of activation in the human brain during an fMRI experiment.
Proceedings of the National Academy of SciencesShrinkage-based similarity metric for cluster analysis of microarray data
41 Citations2003Vera Cherepinsky, Jia‐Wu Feng +2 more
The estimated false positives and false negatives from this study indicate that using the shrinkage metric improves the accuracy of the analysis and compared similarity metrics on a biological example by using the data set from Eisen et al.
Massive computingClustering in Massive Data Sets
39 Citations2002Fionn Murtagh
The visualization of data leads to the consideration of data arrays as images, and theoretical results developed as far back as the 1960s still very often remain topical.
Data Mining and Knowledge DiscoverySampling and Subsampling for Cluster Analysis in Data Mining: With Applications to Sky Survey Data
36 Citations2003David M. Rocke, Jian Dai
The method is quick and reliable and produces classifications comparable to previous work on these data using supervised clustering, and is tested using a sample from the Digitized Palomar Sky Survey data.
International Journal of Offender Therapy and Comparative CriminologyPersonality Typologies of Male Juvenile Offenders Using a Cluster Analysis of the Millon Adolescent Clinical Inventory Introduction
36 Citations2004Tres Stefurak, Georgia B. Calhoun +1 more
The largest group consisted of the reactive depressives, which suggests the importance of considering the role of internalizing problems as a conduit to delinquency in addition to antisocial personality.
TechnometricsClustering Massive Datasets With Application in Software Metrics and Tomography
34 Citations2001Ranjan Maitra
A multistage algorithm that clusters an initial sample, filters out observations that can be reasonably classified by these clusters, and iterates the preceding procedure on the remainder, using the estimated class probabilities and dispersions to classify each observation in the dataset.
Journal of Archaeological SciencePottery production during the Late Jomon period: insights from the chemical analyses of Kasori B pottery
31 Citations2004Mark E. Hall
Journal of Chromatography ATesting of “special base” columns in reversed-phase liquid chromatography
27 Citations2003Katell Le Mapihan, Jérôme Vial +1 more
A methodology for building a chromatographic test aiming at characterizing special base stationary phases and principal component analysis has been combined to hierarchical cluster analysis to reduce drastically the test itself by eliminating redundant information.
Aquatic BotanyA phenetic analysis of Typha in Korea and far east Russia
22 Citations2003Changkyun Kim, Hyunchur Shin +1 more
Principal components analysis and UPGMA analysis showed that individuals of Typha species from Korea and far east Russia form discrete clusters corresponding to four species, and Typha laxmanni was distinguished from T. orientalis by a higher ratio of male and female inflorescence lengths than others.
Journal of Computational and Graphical StatisticsSemi-Streaming Quantization for Remote Sensing Data
15 Citations2003Amy Braverman, Eric J. Fetzer +3 more
This work applies the quantization paradigm from, and algorithms developed in, signal processing to the problem of summarization of very large, remote sensing datasets acquired from NASA's Earth Observing System, and shows that mean squared errors between the final summaries and the original data can be computed.
