login

Scatter/Gather: a cluster-based approach to browsing large document collections

Published 1 January 1992
Douglass R. Cutting, David R. Karger, Jan Pedersen, John W. Tukey
Citations945

TL;DR

This work presents a document browsing technique that employs document clustering as its primary operation, and presents fast (linear time) clustering algorithms which support this interactive browsing paradigm.

Abstract

Document clustering has not been well received as an information retrieval tool. Objections to its use fall into two main categories: first, that clustering is too slow for large corpora (with running time often quadratic in the number of documents); and second, that clustering does not appreciably improve retrieval.

Keywords

Computer Science