Probabilistic document indexing from relevance feedback data
Published 1 December 1989
Norbert Fuhr, Chris Buckley
Citations30
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
Based on the binary independence indexing model, three new concepts for probabilistic document indexing from relevance feedback data are applied:Abstraction from specific terms and documents, Flexibility of the representation, and Probabilistic learning or classification methods for the estimation of the indexing weights making better use of the available relevance information.
Abstract
Based on the binary independence indexing model, we apply three new concepts for probabilistic document indexing from relevance feedback data:
Keywords
Computer Science
Information Processing & ManagementTerm-weighting approaches in automatic text retrieval
9,532 Citations1988Gerard Salton, Chris Buckley
This paper summarizes the insights gained in automatic term weighting, and provides baseline single term indexing models with which other more elaborate content analysis procedures can be compared.
IEEE Transactions on Information TheoryApproximating discrete probability distributions with dependence trees
2,640 Citations1968Chee Lap Chow, C. Liu
It is shown that the procedure derived in this paper yields an approximation of a minimum difference in information when applied to empirical observations from an unknown distribution of tree dependence, and the procedure is the maximum-likelihood estimate of the distribution.
Journal of the ACMOn Relevance, Probabilistic Indexing and Information Retrieval
901 Citations1960M. E. Maron, J. L. Kuhns
The paper suggests an interpretation of the whole library problem as one where the request is considered as a clue on the basis of which the library system makes a concatenated statistical inference in order to provide as an output an ordered list of those documents which most probably satisfy the information needs of the user.
Journal of DocumentationA THEORETICAL BASIS FOR THE USE OF CO‐OCCURRENCE DATA IN INFORMATION RETRIEVAL
465 Citations1977C. J. van Rijsbergen
This paper provides a foundation for a practical way of improving the effectiveness of an automatic retrieval system by measuring the extent of the dependence between index terms and using it to construct a non‐linear weighting function.
Journal of the American Society for Information ScienceA theory of term importance in automatic text analysis
388 Citations1975Gerard Salton, Chul‐Su Yang +1 more
Most existing automatic content analysis and indexing techniques are based on word frequency characteristics applied largely in an ad hoc manner, but terms exhibiting high occurence frequencies in individual documents are often useful for high recall performance, whereas terms with low frequency in the whole collection are useful forhigh precision.
International ACM SIGIR Conference on Research and Development in Information RetrievalProbabilistic models of indexing and searching
313 Citations1980Stephen Robertson, C. J. van Rijsbergen +1 more
There is a considerable body of related work by Salton, Yu and associates on automatic indexing using within-document frequencies of terms.
Communications of the ACMProbabilistic and genetic algorithms in document retrieval
232 Citations1988Michael Gordon
Competing document descriptions are associated with a document and altered over time by a genetic algorithm according to the queries used and relevance judgments made during retrieval.
IEEE Transactions on Pattern Analysis and Machine IntelligenceSynthesizing Statistical Knowledge from Incomplete Mixed-Mode Data
190 Citations1987Andrew K. C. Wong, David Chiu
The proposed method adopts an event-covering approach which covers a subset of statistically relevant outcomes in the outcome space of variable-pairs and can acquire statistical knowledge from incomplete mixed-mode data.
Information Processing & ManagementModels for retrieval with probabilistic indexing
177 Citations1989Norbert Fuhr
Three retrieval models for probabilistic indexing are described along with evaluation results for each, including the binary independence indexing (BII) model, which is a generalized version of the Maron and Kuhns indexing model.
ACM Transactions on Information SystemsOptimum polynomial retrieval functions based on the probability ranking principle
135 Citations1989Norbert Fuhr
This approach is not suited to log-linear probabilistic models and it needs large samples of relevance feedback data for its application, but it can handle very complex representations of documents and requests and it can be easily applied to multivalued relevance scales.
Journal of the American Society for Information ScienceThe effectiveness of a nonsyntactic approach to automatic phrase indexing for document retrieval
132 Citations1989Joel L. Fagan
It is not likely that phrase indexing of this kind will prove to be an important method of enhancing the performance of automatic document indexing and retrieval systems in operational environments, and a general syntactic analysis facility may be required.
Automatic phrase indexing for document retrieval
97 Citations1987Joel L. Fagan
An automatic phrase indexing method based on the term discrimination model is described, and the results of retrieval experiments on five document collections are presented.
A neural network for probabilistic information retrieval
96 Citations1989K. L. Kwok
This paper demonstrates how a neural network may be constructed, together with learning algorithms and modes of operation, that will provide retrieval effectiveness similar to that of the probabilistic indexing and retrieval model based on single terms as document components.
The automatic indexing system AIR/PHYS - from research to applications
60 Citations1988P. Biebricher, Norbert Fuhr +3 more
An appropriate indexing approach and the corresponding structure of the AIR/PHYS system are described, and the conditions of the application as well as problems of further development are discussed.
Journal of the American Society for Information ScienceDocument representation in probabilistic models of information retrieval
51 Citations1981W. Bruce Croft
This article describes how retrieval models which use either independence or dependence assumptions can be extended to include document representatives containing term significance weights and indicates that search strategies based on models modified in this way can further improve the effectiveness of document retrieval systems.
Information Processing & ManagementA probability distribution model for information retrieval
44 Citations1989S. K. M. Wong, Yiyu Yao
A probability distribution model for information retrieval is proposed that not only enhances retrieval effectiveness as demonstrated by experiments, but also provides valuable insight into many fundamental concepts introduced over the years in a variety of retrieval models.
Incorporating syntactic information into a document retrieval strategy
33 Citations1986Alan F. Smeaton
The definition of a retrieval strategy which incorporates parsing of query text and a more “shallow” parsing of document texts, whose retrieval effectiveness is investigated and described are described.
An interpretation of index term weighting schemes based on document components
13 Citations1986K. L. Kwok
It turns out that different choices of document components can lead to different term weighting schemes that have been introduced before and are based on probability considerations; specifically, Edmundson and Wyllys' term significance formula, Sparck Jones' inverse document frequency, and later modified by Croft and Harper into the 'combination match' formula.
Information Processing & ManagementExperiments with document components for indexing and retrieval
12 Citations1988K. L. Kwok, William Kuan
A number of probabilistic similarity measures based on document components are studied, as well as a new method of handling probability estimates involving small sample sizes, and some of the new similarity measures can provide comparable performance to those methods studied by other investigators.
Two learning schemes in information retrieval
9 Citations1988C. Yu, Hidenobu Mizuno
Two methods are given to improve weighting schemes by using relevance information of a set of queries to estimate parameter values of two independence models in information retrieval — the binary independence model and the non-binary independence model.
Optimum polynomial retrieval functions
2 Citations1989Norbert Fuhr
In contrast to other probabilistic models, this approach yields estimates of the actual probabilities, it can handle very complex representations of documents and requests, and it can be easily applied to multi-valued relevance scales.
