login

Extension of the mixture of factor analyzers model to incorporate the multivariate t-distribution

Computational Statistics & Data AnalysisPublished 11 October 2006
Geoffrey J. McLachlan, Richard Bean, L. Ben-Tovim Jones
Citations151
SJR quartileQ1
SJR score0.89
SNIP1.38

TL;DR

An EM-based algorithm is developed for the fitting of mixtures of t-factor analyzers and its application is demonstrated in the clustering of some microarray gene-expression data.

Abstract

Mixtures of factor analyzers enable model-based density estimation to be undertaken for high-dimensional data, where the number of observations n is small relative to their dimension p. However, this approach is sensitive to outliers as it is based on a mixture model in which the multivariate normal family of distributions is assumed for the component error and factor distributions. An extension to mixtures of t-factor analyzers is considered, whereby the multivariate t-family is adopted for the component error and factor distributions. An EM-based algorithm is developed for the fitting of mixtures of t-factor analyzers. Its application is demonstrated in the clustering of some microarray gene-expression data.

Keywords

Computer ScienceMathematicsBiochemistry, Genetics and Molecular Biology