Model-Based Clustering for Image Segmentation and Large Datasets via Sampling
Journal of ClassificationPublished 1 September 2004
Ron Wehrens, L.M.C. Buydens, Chris Fraley, Adrian E. Raftery
Citations78
SJR quartileQ1
SJR score0.71
SNIP1.33
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
These experiments suggest that a stable method with better performance can be obtained with two straightforward modifications to the simple sampling method: several tentative models are identified from the sample instead of just one, and several EM steps are used rather than just one E step to classify the full data set.
Abstract
\n Contains fulltext :\n 60419.pdf (Publisher’s version ) (Closed access)\n
Keywords
Computer Science
Journal of the Royal Statistical Society Series B (Statistical Methodology)Maximum Likelihood from Incomplete Data Via the <i>EM</i> Algorithm
49,657 Citations1977A. P. Dempster, N. M. Laird +1 more
Choice Reviews OnlineFinding groups in data: an introduction to cluster analysis
10,614 Citations1991
Journal of Computational and Graphical StatisticsR: A Language for Data Analysis and Graphics
10,424 Citations1996Ross Ihaka, Robert Gentleman
The experience designing and implementing a statistical computing language that provides advantages in the areas of portability, computational efficiency, memory management, and scoping is discussed.
Wiley series in probability and statisticsFinite Mixture Models
7,410 Citations2018Geoffrey J. McLachlan, David Peel +1 more
The aim of this article is to provide an up-to-date account of the theory and methodological developments underlying the applications of finite mixture models.
Journal of the American Statistical AssociationObjective Criteria for the Evaluation of Clustering Methods
5,815 Citations1971Telmo Menezes
TechnometricsThe EM Algorithm and Extensions
5,108 Citations1998Debashis Kushary, Geoffrey J. McLachlan +1 more
Journal of the American Statistical AssociationModel-Based Clustering, Discriminant Analysis, and Density Estimation
4,282 Citations2002Chris Fraley, Adrian E. Raftery
This work reviews a general methodology for model-based clustering that provides a principled statistical approach to important practical questions that arise in cluster analysis, such as how many clusters are there, which clustering method should be used, and how should outliers be handled.
The Computer JournalHow Many Clusters? Which Clustering Method? Answers Via Model-Based Cluster Analysis
2,519 Citations1998Chris Fraley
The problems of determining the number of clusters and the clustering method are solved simultaneously by choosing the best model, and the EM result provides a measure of uncertainty about the associated classification of each data point.
BiometricsModel-Based Gaussian and Non-Gaussian Clustering
2,350 Citations1993Jeffrey D. Banfield, Adrian E. Raftery
The classification maximum likelihood approach is sufficiently general to encompass many current clustering algorithms, including those based on the sum of squares criterion and on the criterion of Friedman and Rubin (1967), but it is restricted to Gaussian distributions and it does not allow for noise.
A View of the Em Algorithm that Justifies Incremental, Sparse, and other Variants
2,187 Citations1998Radford M. Neal, Geoffrey E. Hinton
An incremental variant of the EM algorithm in which the distribution for only one of the unobserved variables is recalculated in each E step is shown empirically to give faster convergence in a mixture estimation problem.
Queensland's institutional digital repository (The University of Queensland)Mixture models : inference and applications to clustering
2,074 Citations1988Geoffrey J. McLachlan, K. E. Basford
Journal of the American Statistical AssociationComputing Bayes Factors by Combining Simulation and Asymptotic Approximations
1,967 Citations1997Thomas J. DiCiccio, Robert E. Kass +2 more
It is found that a simulated version of Laplace's method, with local volume correction, furnishes an accurate approximation that is especially useful when likelihood function evaluations are costly.
BioinformaticsModel-based clustering and data transformations for gene expression data
888 Citations2001Ka Yee Yeung, Chris Fraley +3 more
Pattern RecognitionGaussian parsimonious clustering models
842 Citations1995Gilles Celeux, Gérard Govaert
Methods of optimization to derive the maximum likelihood estimates as well as the practical usefulness of these models are discussed and an application on stellar data which dramatically illustrated the relevance of allowing clusters to have different volumes is illustrated.
Journal of ClassificationMCLUST: Software for Model-Based Cluster Analysis
500 Citations1999Chris Fraley, Adrian E. Raftery
MCLUST is a software package for cluster analysis written in Fortran and interfaced to the S PLUS commercial software package and includes functions that combine hierarchical clustering EM and the Bayesian Information Criterion BIC in a comprehensive clustering strategy.
Journal of the American Statistical AssociationDetecting Features in Spatial Point Processes with Clutter via Model-Based Clustering
415 Citations1998Abhijit Dasgupta, Adrian E. Raftery
This work uses model-based clustering based on a mixture model for the problem of detecting features in spatial point processes when there is substantial clutter, in which features are assumed to generate points according to highly linear multivariate normal densities, and the clutter arises according to a spatial Poisson process.
Journal of ClassificationEnhanced Model-Based Clustering, Density Estimation, and Discriminant Analysis Software: MCLUST
334 Citations2003Chris Fraley, Adrian E. Raftery
MCLUST is a software package for model-based clustering, density estimation and discriminant analysis interfaced to the S-PLUS commercial software and the R language that implements parameterized Gaussian hierarchical clustering algorithms and the EM algorithm for parameterizedGaussian mixture models with the possible addition of a Poisson noise term.
The Astrophysical JournalThree Types of Gamma‐Ray Bursts
296 Citations1998Soma Mukherjee, Eric D. Feigelson +4 more
SIAM Journal on Scientific ComputingAlgorithms for Model-Based Gaussian Hierarchical Clustering
268 Citations1998Chris Fraley
It is shown how the structure of the Gaussian model can be exploited to yield efficient algorithms for agglomerative hierarchical clustering.
Statistics and ComputingInference in model-based cluster analysis
201 Citations1997Halima Bensmail, Gilles Celeux +2 more
This work proposes a new approach to cluster analysis which consists of exact Bayesian inference via Gibbs sampling, and the calculation of Bayes factors from the output using the Laplace–Metropolis estimator, which works well in several real and simulated examples.
Pattern Recognition LettersLinear flaw detection in woven textiles using model-based clustering
153 Citations1997Jonathan Campbell, Chris Fraley +2 more
This approach detects a linear pattern in preprocessed images via model-based clustering using an approximate Bayes factor which provides a criterion for assessing the evidence for the presence of a defect.
The EMMIX software for the fitting of mixtures of normal and t-components
143 Citations1999Geoffrey J. McLachlan, David Peel +2 more
IEEE Transactions on Pattern Analysis and Machine IntelligenceFinding curvilinear features in spatial point patterns: principal curve clustering with noise
129 Citations2000Derek C. Stanford, Adrian E. Raftery
The algorithm for principal curve clustering is in two steps: the first is hierarchical and agglomerative (HPCC) and the second consists of iterative relocation based on the classification EM algorithm (CEM-PCC), which is used to combine potential feature clusters and refines the results and deals with background noise.
MCLUST: Software for Model-Based Clustering, Density Estimation and Discriminant Analysis
127 Citations2002Chris Fraley, Adrian E. Raftery
MCLUST is a software package for model-based clustering, density estimation and discriminant analysis interfaced to the S-PLUS commercial software that implements parameterized Gaussian hierarchical clustering algorithms and the EM algorithm for parameterizedGaussian mixture models with the possible addition of a Poisson noise term.
Handbook of statistics9 The classification and mixture maximum likelihood approaches to cluster analysis
103 Citations1982Geoffrey J. McLachlan
The chapter discusses the efficiency of the mixture approach for k = 2 normal subpopulations, contrasting the asymptotic theory with small sample results available from simulation.
Journal of Computational and Graphical StatisticsIncremental Model-Based Clustering for Large Datasets With Small Clusters
81 Citations2005Chris Fraley, Adrian E. Raftery +1 more
An incremental approach for data that can be processed as a whole in memory is proposed, which is relatively efficient computationally and has the ability to find small clusters in large datasets.
Journal of the American Statistical AssociationNearest-Neighbor Variance Estimation (NNVE)
74 Citations2002Naisyin Wang, Adrian E. Raftery
The “outlyingness” of a data point is measured by the standardized distance between the point and its Kth nearest neighbor, which reduces the problem of finding the robustness weights to a one-dimensional problem and may be useful in high-dimensional problems, such as those encountered in data mining.
International Journal of Imaging Systems and TechnologyModel-based methods for textile fault detection
60 Citations1999Jonathan Campbell, Chris Fraley +3 more
A model‐based clustering method that can be employed to aggregate perceptual groupings of point and local detections and a discrete Fourier transform‐based texture analysis technique that is highly effective for woven textiles in discriminating subtle flaw patterns from the pronounced background of repetitive weaving pattern and random clutter are described.
IEEE Transactions on Pattern Analysis and Machine IntelligenceApproximate Bayes factors for image segmentation: the Pseudolikelihood Information Criterion (PLIC)
59 Citations2002Derek C. Stanford, Adrian E. Raftery
A method for choosing the number of colors or true gray levels in an image; this allows fully automatic segmentation of images and discusses a simpler approximation, MMIC (Marginal Mixture Information Criterion), which is based only on the marginal distribution of pixel values.
Journal of Computational and Graphical StatisticsHierarchical Model-Based Clustering for Large Datasets
44 Citations2001Christian Posse
This article proposes to start the hierarchical agglomeration from an efficient classification of the data in many classes rather than from the usual set of singleton clusters, and develops graphical tools that assess the presence of clusters in the data and uncover observations difficult to classify.
Statistics and ComputingOn the choice of the number of blocks with the incremental EM algorithm for the fitting of normal mixtures
39 Citations2003Shu‐Kay Ng, Geoffrey J. McLachlan
A simple rule is proposed for choosing the number of blocks with the IEM algorithm in the extreme case of one observation per block, which provides efficient updating formulas, which avoid the direct calculation of the inverses and determinants of the component-covariance matrices.
Hierarchical model-based clustering of large datasets through fractionation and refractionation
39 Citations2002Jeremy Tantrum, Alejandro Murua +1 more
An adaptation of Fractionation to model-based clustering, originally conceived by Cutting, Karger, Pedersen and Tukey, is described and a further extension, called Refractionation, leads to a procedure that can be successful even in the difficult situation where there are large numbers of small groups.
TechnometricsClustering Massive Datasets With Application in Software Metrics and Tomography
34 Citations2001Ranjan Maitra
A multistage algorithm that clusters an initial sample, filters out observations that can be reasonably classified by these clusters, and iterates the preceding procedure on the remainder, using the estimated class probabilities and dispersions to classify each observation in the dataset.
Journal of ChemometricsMixture modelling of medical magnetic resonance data
26 Citations2002Ron Wehrens, Arjan W. Simonetti +1 more
Queensland's institutional digital repository (The University of Queensland)Clustering of magnetic resonance images
14 Citations1996Geoffrey J. McLachlan, Shu‐Kay Ng +2 more
A statistical model-based approach to the clustering of multispectral magnetic resonance (MR) images using an approximation to the E-step based on the ICM algorithm is considered.
