login

Comparing and Unifying Search-Based and Similarity-Based Approaches to Semi-Supervised Clustering

Published 1 January 2003
Sugato Basu, Mikhail Bilenko, Raymond J. Mooney
Citations81

TL;DR

A unifled approach based on the K-Means clustering algorithm that incorporates both searchbased and similarity-based techniques, and demonstrates that the combined approach generally produces better clusters than either of the individual approaches.

Abstract

Semi-supervised clustering employs a small amount of labeled data to aid unsupervised learning. Previous work in the area has employed one of two approaches: 1) Search-based methods that utilize supervised data to guide the search for the best clustering, and 2) Similarity-based methods that use supervised data to adapt the underlying similarity metric used by the clustering algorithm. This paper presents a unified approach based on the K-Means clustering algorithm that incorporates both of these techniques. Experimental results demonstrate that the combined approach generally produces better clusters than either of the individual approaches.

Keywords

Computer ScienceDecision Sciences