login

Restructuring Sparse High Dimensional Data for Effective Retrieval

DSpace@MIT (Massachusetts Institute of Technology)Published 1 December 1998Open access
Charles L. Isbell, Paul Viola
Citations51
View PDF

TL;DR

This work proposes an alternative and novel technique that produces sparse representations constructed from sets of highly-related words that significantly improves retrieval performance, is efficient to compute and shares properties with the optimal linear projection operator and the independent components of documents.

Abstract

The task in text retrieval is to find the subset of a collection of documents relevant to a user’s information request, usually expressed as a set of words. Classically, documents and queries are represented as vectors of word counts. In its simplest form, relevance is defined to be the dot product between a document and a query vector–a measure of the number of common terms. A central difficulty in text retrieval is that the presence or absence of a word is not sufficient to determine relevance to a query. Linear dimen-sionality reduction has been proposed as a technique for extracting underlying structure from the document collection. In some domains (such as vision) dimensionality re-duction reduces computational complexity. In text retrieval it is more often used to improve retrieval performance. We propose an alternative and novel technique that pro-duces sparse representations constructed from sets of highly-related words. Documents and queries are represented by their distance to these sets. and relevance is measured by the number of common clusters. This technique significantly improves retrieval perfor-mance, is efficient to compute and shares properties with the optimal linear projection operator and the independent components of documents. Copyright c©Massachusetts Institute of Technology, 1998. This publication can be retrieved by anonymous ftp at

Keywords

Computer Science