login

On query formulation in information retrieval

Published 1 January 1981
Harry Wu
Citations9

TL;DR

Experimental evidence indicates that relevance feedback is a most promising strategy for query term weights and the advantages and disadvantages of the Boolean and vector representations of queries are demonstrated.

Abstract

Various schemes of query formulation in information retrieval are analyzed. The frequency characteristics of terms in the documents of a collection have long been used as indicators of term importance for indexing purposes. In particular, very rare or very frequent terms are normally believed to be less effective than the medium frequency terms. Accordingly, it has been suggested that query term weights should first increase and then decrease as the document frequencies of the query terms increase. Three term weighting models (utility weights, precision weights, and discrimination value weights) are studied and they all substantiate the suggested behavior. A fourth scheme, the inverse document frequency system, is shown to be a close approximation of the precision system under certain circumstances. Two of the three models require relevance information. The precision and utility weights are based on the occurrence characteristics of the terms in the relevant, as opposed to the nonrelevant, documents in the collection. Methods are suggested for estimating the relevant properties of the terms based on the overall occurrence characteristics in the collection. Empirical evaluation results are shown comparing the weighting systems using the term relevance properties with the more conventional frequency-based methodologies. Experimental evidence indicates that relevance feedback is a most promising strategy. The last part of the thesis demonstrates the advantages and disadvantages of the Boolean and vector representations of queries. A more general framework which encompasses both the fuzzy set approach and the vector space approach is outlined. The new scheme is derived from the concept of p-norms. It is shown that the fuzzy set approach is the extreme case when p = (INFIN) while the vector space approach represents the other extreme case when p = 1.

Keywords

Computer Science