A study of standardization of variables in cluster analysis
Journal of ClassificationPublished 1 September 1988
Glenn W. Milligan, Martha C. Cooper
Citations873
SJR quartileQ1
SJR score0.71
SNIP1.33
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
The present simulation study examined the standardization problem and found that those approaches which standardize by division by the range of the variable gave consistently superior recovery of the underlying cluster structure.
Abstract
Standard scores, Cluster analysis,
Keywords
Computer ScienceMathematicsAgricultural and Biological Sciences
Biometry: The Principles and Practice of Statistics in Biological Research
21,134 Citations1969Robert R. Sokal, F. James Rohlf
Cluster Analysis
9,195 Citations1974Barry J. Everitt, Sabine Landau +1 more
Journal of ClassificationComparing partitions
7,714 Citations1985Lawrence J. Hubert, Phipps Arabie
A measure based on the comparison of object triples having the advantage of a probabilistic interpretation in addition to being corrected for chance is proposed and bounded between ±1.5 and ±2.5.
BiometricsA General Coefficient of Similarity and Some of Its Properties
5,240 Citations1971J. C. Gower
A general coefficient measuring the similarity between two sampling units is defined and the matrix of similarities between all pairs of sample units is shown to be positive semidefinite.
PsychometrikaHierarchical Clustering Schemes
4,891 Citations1967S. C. Johnson
A useful correspondence is developed between any hierarchical system of such clusters, and a particular type of distance measure, that gives rise to two methods of clustering that are computationally rapid and invariant under monotonic transformations of the data.
The American StatisticianRank Transformations as a Bridge between Parametric and Nonparametric Statistics
3,547 Citations1981W. J. Conover, Ronald L. Iman
PsychometrikaAn Examination of the Effect of Six Types of Error Perturbation on Fifteen Clustering Algorithms
1,170 Citations1980Glenn W. Milligan
An evaluation of several clustering methods indicated that the hierarchical methods were differentially sensitive to the type of error perturbation and two alternative starting procedures for the nonhierarchical methods produced greatly enhanced cluster recovery and were found to be robust to all of the types of error examined.
Applied Psychological MeasurementMethodology Review: Clustering Methods
651 Citations1987Glenn W. Milligan, Martha C. Cooper
A review of clustering methodology is presented, with emphasis on algorithm performance and the re sulting implications for applied research, and two sets of recommendations are offered.
Multivariate Behavioral ResearchA Study of the Comparability of External Criteria for Hierarchical Cluster Analysis
488 Citations1986Glenn W. Milligan, Martha C. Cooper
The results of the study indicated that the Hubert and Arabie adjusted Rank index was best suited to the task of comparison across hierarchy levels.
Advances in computersClustering Methodologies in Exploratory Data Analysis
287 Citations1980Richard C. Dubes, Anil Jain
The chapter focuses on the four operations highlighted by reviewing techniques for assessing the tendency of the data to cluster, performing the clustering itself, and evaluating the validity of the results, and introduces the concept of intrinsic dimensionality that helps determine an appropriate number of factors for representing data.
Multivariate Behavioral ResearchA Review Of Monte Carlo Tests Of Cluster Analysis
240 Citations1981Glenn W. Milligan
A review of Monte Carlo validation studies of clustering algorithms indicates that other algorithms may provide better recovery under a variety of conditions than Ward's minimum variance hierarchical method.
Systematic ZoologyDistance as a Measure of Taxonomic Similarity
209 Citations1961Robert R. Sokal
Sudden developments in methods for quantifying the classificatory process in systematics appear to stem from a growing dissatisfaction with the arbitrariness and subjectivity of the customary taxonomic procedure and would also appear to.
Journal of the American Statistical AssociationPower Differences between Pairwise Multiple Comparisons
166 Citations1978Philip H. Ramsey
Proceedings of the Zoological Society of LondonAN ANALYSIS OF THE TAXONOMIST'S JUDGMENT OF AFFINITY
162 Citations1958A. J. Cain, G. Ainsworth Harrison
A procedure is worked out which makes precise the judgment of affinity, and enables us to obtain a numerical value for the mean character difference between any two forms.
Multivariate Behavioral ResearchMixture Model Tests Of Hierarchical Clustering Algorithms: The Problem Of Classifying Everybody
161 Citations1979Craig Edelbrock
A subset of high accuracy algorithms, including single, average, and centroid linkage using correlation, and Ward's minimum variance technique, was identified and all of the algorithms were significantly more accurate than a random linkage algorithm, and accuracy was inversely related to coverage.
PsychometrikaAn Algorithm for Generating Artificial Test Clusters
155 Citations1985Glenn W. Milligan
An algorithm for generating artificial data sets which contain distinct nonoverlapping clusters is presented, useful for generating test data sets for Monte Carlo validation research conducted on clustering methods or statistics.
Multivariate Behavioral ResearchON THE METHODS AND THEORY OF CLUSTERING
142 Citations1969Joseph L. Fleiss, Joseph Zubin
The key defect in almost all clustering procedures seems to be the absence of a statistical model, and the suggestion is made that the clustering problem be stated as a mixture problem.
Multivariate Behavioral ResearchMonte Carlo Tests of the Accuracy of Cluster Analysis Algorithms: A Comparison of Hierarchical and Nonhierarchical Methods
125 Citations1985Dieter Scheibler, Wolfgang Schneider
The results confirmed the findings of previous Monte Carlo studies on clustering procedures in that accuracy was inversely related to coverage, and that algorithms using correlation as the similarity measure were significantly more accurate than those using Euclidean distances.
Pattern RecognitionMonte Carlo comparisons of selected clustering procedures
100 Citations1980Charles K. Bayne, John J. Beauchamp +2 more
Monte Carlo methods were used to estimate the percent misclassification of 13 clustering methods for six types of parameterizations of two bivariate normal populations and the overall poorest methods were judged to be nearest neighbor and maximum likelihood.
Management ScienceMeasurement Problems in Cluster Analysis
98 Citations1967Donald G. Morrison
The first part of this paper will review and modify the cluster analysis procedure presented by Green, Frank and Robinson and raise some very fundamental questions with respect to cluster analysis in particular and multivariate statistics in general.
Journal of ClassificationOptimal variable weighting for hierarchical clustering: An alternating least-squares algorithm
55 Citations1985Geert De Soete, Wayne S. DeSarbo +1 more
A new methodology which simultaneously estimates in a least-squares fashion both an ultrametric tree and respective variable weightings for profile data that have been converted into (weighted) Euclidean distances is presented.
NatureAn Objective Method of Weighting in Similarity Analysis
31 Citations1964W. T. Williams, M. B. Dale +1 more
BiometricsStandardization of Measures Prior to Cluster Analysis
25 Citations1979Anne M. Stoddard
A procedure for scaling measurements using a reference individual as a standard of comparison is presented, which appears to remove extraneous variability while retaining the information necessary for classification in cluster analysis.
Sociological Methods & ResearchIssues in Multivariate Cluster Analysis
15 Citations1985Robert L. Kaufman
Using a Monte Carlo simulation, the research addresses two key questions about the accuracy of cluster analysis in reproducing a known true cluster model and indicates that using principal components analysis is superior to not using it and that the choice of how to utilize the principal components results may be critical.
PsychometrikaThe Equivalence of Three Statistical Packages for Performing Hierarchical Cluster Analysis
14 Citations1977Roger K. Blashfield
Three different software programs which contain hierarchical agglomerative cluster analysis procedures were shown to generate different solutions on the same data set using apparently the same options, the basis for the differences was the formulae used to calculate Euclidean distance.
NatureThe Peculiarity Index, a New Function for Use in Numerical Taxonomy
11 Citations1965A. V. Hall
A function was derived which may be referred to as a ‘Peculiarity Index’ to show the relative proportions of unusual features in taxa and it was found that for some characters one state was much rarer than the other.
Biometrical JournalWeighted Standardization—A General Data Transformation Method Proceeding Classification Procedures
10 Citations1986Johann Hohenegger
During preparatory steps of data for automatic classification routines, the amount of information contained by the character distribution is reduced by standardization of the character values but can be regained through special weighting schemes of standardized character values.
