login

FREM

Published 4 November 2002
Carlos Ordońẽz, Edward Omiecinski
Citations69

TL;DR

An improved EM algorithm to cluster large data sets having high dimensionality, noise and zero variance problems is presented and is compared against the standard EM algorithm and the On-Line EM algorithm.

Abstract

Clustering is a fundamental Data Mining technique. This article presents an improved EM algorithm to cluster large data sets having high dimensionality, noise and zero variance problems. The algorithm incorporates improvements to increase the quality of solutions and speed. In general the algorithm can find a good clustering solution in 3 scans over the data set. Alternatively, it can be run until it converges. The algorithm has a few parameters that are easy to set and have defaults for most cases. The proposed algorithm is compared against the standard EM algorithm and the On-Line EM algorithm.

Keywords

Computer Science