Probability models for information retrieval based on divergence from randomness
Published 1 January 2003
Giambattista Amati
Citations144
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
This thesis devises a novel methodology based on probability theory, suitable for the construction of term-weighting models of Information Retrieval, and shows that even language modelling approach can be exploited to assign term-frequency normalization to the models of divergence from randomness.
Abstract
Available from British Library Document Supply Centre- DSC:DXN065251 / BLDSC - British Library Document Supply Centre
Keywords
Computer Science
Bell System Technical JournalA Mathematical Theory of Communication
80,675 Citations1948Claude E. Shannon
Journal of the Royal Statistical Society Series B (Statistical Methodology)Maximum Likelihood from Incomplete Data Via the <i>EM</i> Algorithm
49,657 Citations1977A. P. Dempster, N. M. Laird +1 more
Journal of the Franklin InstituteAn introduction to probability theory and its applications
29,966 Citations1958
Information Processing & ManagementTerm-weighting approaches in automatic text retrieval
9,532 Citations1988Gerard Salton, Chris Buckley
This paper summarizes the insights gained in automatic term weighting, and provides baseline single term indexing models with which other more elaborate content analysis procedures can be compared.
TechnometricsThe EM Algorithm and Extensions
5,108 Citations1998Debashis Kushary, Geoffrey J. McLachlan +1 more
Journal of DocumentationA STATISTICAL INTERPRETATION OF TERM SPECIFICITY AND ITS APPLICATION IN RETRIEVAL
4,442 Citations1972Karen Spärck Jones
It is argued that terms should be weighted according to collection frequency, so that matches on less frequent, more specific, terms are of greater value than matches on frequent terms.
Journal of the American Statistical AssociationStatistical Analysis of Finite Mixture Distributions.
2,940 Citations1987Bruce G. Lindsay, D. M. Titterington +2 more
Information Retrieval: Data Structures and Algorithms
2,428 Citations1992William B. Frakes, Ricardo Baeza‐Yates
For programmers and students interested in parsing text, automated indexing, its the first collection in book form of the basic data structures and algorithms that are critical to the storage and retrieval of documents.
Text REtrieval ConferenceOkapi at TREC
2,229 Citations1994Stephen Robertson, Steve Walker +3 more
Much of the work involved investigating plausible methods of applying Okapi-style weighting to phrases, and expansion using terms from the top documents retrieved by a pilot search on topic terms was used.
Journal of the American Society for Information ScienceRelevance weighting of search terms
2,068 Citations1976Stephen Robertson, Karen Spärck Jones
This paper examines statistical techniques for exploiting relevance information to weight search terms using information about the distribution of index terms in documents in general and shows that specific weighted search methods are implied by a general probabilistic theory of retrieval.
Information and ControlA formal theory of inductive inference. Part II
1,786 Citations1964Ray J. Solomonoff
Four ostensibly different theoretical models of induction are presented, in which the problem dealt with is the extrapolation of a very long sequence of symbols—presumably containing all of the information to be used in the induction.
ACM SIGIR ForumA Study of Smoothing Methods for Language Models Applied to Ad Hoc Information Retrieval
1,571 Citations2017ChengXiang Zhai, John Lafferty
This paper examines the sensitivity of retrieval performance to the smoothing parameters and compares several popular smoothing methods on different test collection.
Journal of the American Society for Information ScienceImproving retrieval performance by relevance feedback
1,504 Citations1990Gerard Salton, Chris Buckley
Relevance feedback is an automatic process, introduced over 20 years ago, designed to produce query formulations following an initial retrieval operation to demonstrate the effectiveness of the various methods.
ACM SIGIR ForumQuary Expansion Using Local and Global Document Analysis
1,269 Citations2017Jinxi Xu, W. Bruce Croft
It is shown that, although global analysis haa some advantages, local analysia is generally more effective than global techniques and using global analysis techniques.
Journal of DocumentationTHE PROBABILITY RANKING PRINCIPLE IN IR
1,071 Citations1977Stephen Robertson
It is shown that the principle that documents should be ranked in order of the probability of relevance or usefulness can be justified under certain assumptions, but that in cases where these assumptions do not hold, the principle is not valid.
IBM Journal of Research and DevelopmentA Statistical Approach to Mechanized Encoding and Searching of Literary Information
1,059 Citations1957H. P. Luhn
The problem of literature searching by machines still presents major difficulties and a statistical approach to this problem will be outlined and the various steps of a system based on this approach will be described.
Journal of the ACMOn Relevance, Probabilistic Indexing and Information Retrieval
901 Citations1960M. E. Maron, J. L. Kuhns
The paper suggests an interpretation of the whole library problem as one where the request is considered as a clue on the basis of which the library system makes a concatenated statistical inference in order to provide as an output an ordered list of those documents which most probably satisfy the information needs of the user.
ACM Transactions on Information SystemsProbabilistic models of information retrieval based on measuring the divergence from randomness
883 Citations2002Gianni Amati, Cornelis J. van Rijsbergen
A framework for deriving probabilistic models of Information Retrieval using term-weighting models obtained in the language model approach by measuring the divergence of the actual term distribution from that obtained under a random process is introduced.
American Journal of PhysicsTransmission of Information: A Statistical Theory of Communications
866 Citations1961Robert M. Fano, David Hawkins
ACM SIGIR ForumPivoted Document Length Normalization
863 Citations2017Amit Singhal, Chris Buckley +1 more
Pivoted normalization is presented, a technique that can be used to modify any normalization function thereby reducing the gap between the relevance and the retrieval probabilities, and two new normalization functions--pivoted unique normalization and piuotert byte size normalization are presented.
ACM SIGIR ForumDocument Language Models, Query Models, and Risk Minimization for Information Retrieval
774 Citations2017John Lafferty, ChengXiang Zhai
A framework for information retrieval that combines document models and query models using a probabilistic ranking function based on Bayesian decision theory is presented and an operational retrieval model that extends recent developments in the language modeling approach to information retrieval is suggested.
ACM Transactions on Information SystemsEvaluation of an inference network-based retrieval model
577 Citations1991Howard R. Turtle, W. Bruce Croft
Network representations show promise as mechanisms for inferring probable relationships between documents and queries and have been used in information retrieval since at least the early 1960s.
Improving automatic query expansion
555 Citations1998Mandar Mitra, Amit Singhal +1 more
Experimental results show that refining the set of documents used in query expansion often prevents the query drift caused by blind expansion and yields substantial improvements in retrieval effectiveness, both in terms of average precision and precision in the top twenty documents.
Journal of the ACMAutomatic Indexing: An Experimental Inquiry
552 Citations1961M. E. Maron
The design, execution and evaluation of a modest experimental study aimed at testing empirically one statistical technique for automatic indexed documents according to their subject content are described.
ACM Transactions on Information SystemsImproving the effectiveness of information retrieval with local context analysis
533 Citations2000Jinxi Xu, W. Bruce Croft
A new technique is proposed, called local context analysis, which selects expansion terms based on cooccurrence with the query terms within the top-ranked documents.
Overview of the TREC 2006.
469 Citations2006Ellen M. Voorhees
The fourteenth Text REtrieval Conference, TREC 2005, was held at the National Institute of Standards and Technology (NIST) 15 to 18 November 2005 and had 117 participating groups from 23 different countries.
Journal of DocumentationA THEORETICAL BASIS FOR THE USE OF CO‐OCCURRENCE DATA IN INFORMATION RETRIEVAL
465 Citations1977C. J. van Rijsbergen
This paper provides a foundation for a practical way of improving the effectiveness of an automatic retrieval system by measuring the extent of the dependence between index terms and using it to construct a non‐linear weighting function.
Journal of DocumentationUSING PROBABILISTIC MODELS OF DOCUMENT RETRIEVAL WITHOUT RELEVANCE INFORMATION
436 Citations1979W. Bruce Croft, David J. Harper
This paper considers the situation where no relevance information is available, that is, at the start of the search, based on a probabilistic model, and proposes strategies for the initial search and an intermediate search.
ACM SIGIR ForumInference Networks for Document Retrieval
425 Citations2017Howard R. Turtle, W. Bruce Croft
The use of inference networks to support document retrieval and a network-basead retrieval model is described and compared to conventional probabilistic and Boolean models.
Relevance feedback revisited
405 Citations1992Donna Harman
These experiments, using the Cranfield 1400 collection, showed the importance of query expansion in addition to query reweighting, and showed that adding as few as 20 well-selected terms could result in performance improvements of over 100%.
IEEE Transactions on Information TheoryComplexity-based induction systems: Comparisons and convergence theorems
404 Citations1978Ray J. Solomonoff
Levin has shown that if tilde{P}'_{M}(x) is an unnormalized form of this measure, and P( x) is any computable probability measure on strings, x, then \tilde{M}'_M}\geqCP (x) where C is a constant independent of x .
ACM Transactions on Information SystemsAn information-theoretic approach to automatic query expansion
366 Citations2001Claudio Carpineto, Renato De Mori +2 more
This work presents a computationally simple and theoretically justified method for assigning scores to candidate expansion terms within Rocchio's framework for query reweigthing, and discusses the effect on retrieval effectiveness of the main parameters involved in automatic query expansion.
The Economic JournalThe Theory of Income Distribution.
337 Citations1974Martin Bronfenbrenner, Harry G. Johnson
Heavy-tailed probability distributions in the World Wide Web
328 Citations1998Mark Crovella, Murad S. Taqqu +1 more
Evidence is presented that a number of le size distributions in the Web exhibit heavy tails, including les requested by users, les transmitted through the network, transmission durations of les, and les stored on servers, that are primarily determined by the distribution of les available on the Web.
International ACM SIGIR Conference on Research and Development in Information RetrievalProbabilistic models of indexing and searching
313 Citations1980Stephen Robertson, C. J. van Rijsbergen +1 more
There is a considerable body of related work by Salton, Yu and associates on automatic indexing using within-document frequencies of terms.
Information Processing & ManagementOverview of the Second Text Retrieval Conference (TREC-2)
310 Citations1995Donna Harman
This conference, co-sponsored by ARPA and NIST, brought together information retrieval researchers to discuss their system results on the new TIPSTER test collection, and represented a breakthrough in cross-system evaluation in information retrieval.
The Computer JournalProbabilistic Models in Information Retrieval
308 Citations1992Norbert Fuhr
An introduction and survey over probabilistic information retrieval (IR) is given: the probability-ranking principle shows that optimum retrieval quality can be achieved under certain assumptions; a conceptual model for IR along with the corresponding event space clarify the interpretation of the Probabilistic parameters involved.
Information Storage and RetrievalA definition of relevance for information retrieval
306 Citations1971William S. Cooper
A definition of what it means to say that a piece of stored information is “relevant” to the information need of a retrieval system user is proposed and defended and explicates relevance in terms of logical implication.
Overview of the Eighth Text REtrieval Conference (TREC-8).
305 Citations1999Ellen M. Voorhees, Donna Harman
Language and Representation in Information Retrieval
260 Citations1990David C. Blair
This work has shown that language and Representation are the central problem in Information Retrieval and the nature of scientific theory, and the principal formal models used in information retrieval are language and representation.
Journal of the ACMLocal Feedback in Full-Text Retrieval Systems
234 Citations1977Rony Attar, Aviezri S. Fraenkel
Local clustering is practical also for large databases and appears to improve overall performance, especially if metrical constraints and weighting by proximity are embedded m the local feedback.
American Journal of PhysicsThe Algebra of Probable Inference
230 Citations1963R. T. Cox, E. T. Jaynes
Modeling score distributions for combining the outputs of search engines
208 Citations2001R. Manmatha, T.M. Rath +1 more
It is shown empirically that the score distributions of a number of text search engines on a per query basis may be fitted using an exponential distribution for the set of non-relevant documents and a normal distribution forThe set of relevant documents.
Information Processing & ManagementDocument length normalization
202 Citations1996Amit Singhal, Gerard Salton +2 more
A modified technique is presented that attempts to match the likelihood of retrieving a document of a certain length to thelihood of documents of that length being judged relevant, and it is shown that this technique yields significant improvements in retrieval effectiveness.
Journal of the American Society for Information ScienceA probabilistic approach to automatic keyword indexing. Part I. On the Distribution of Specialty Words in a Technical Literature
198 Citations1975Stephen P. Harter
A mixture of two Poisson distributions is examined in detail as a model of specialty word distribution and a measure intended to identify specialty words, consistent with the 2-Poisson model, is proposed and evaluated.
ACM Computing Surveys“Is this document relevant?…probably”
197 Citations1998Fábio Crestani, Mounia Lalmas +2 more
The basic concepts of probabilistic approaches to information retrieval are outlined and the principles and assumptions upon which the approaches are based are presented.
Two-stage language models for information retrieval
193 Citations2002ChengXiang Zhai, John Lafferty
Evaluation on five different databases and four types of queries indicates that the two-stage smoothing method with the proposed parameter estimation methods consistently gives retrieval performance that is close to---or better than---the best results achieved using a single smoothed method and exhaustive parameter search on the test data.
ACM Transactions on Information SystemsOn modeling information retrieval with probabilistic inference
185 Citations1995S. K. M. Wong, Yiyu Yao
This article examines and extends the logical models of information retrieval in the context of probability theory, and the fundamental notions of term weights and relevance are given probabilistic interpretations.
Information Processing & ManagementModels for retrieval with probabilistic indexing
177 Citations1989Norbert Fuhr
Three retrieval models for probabilistic indexing are described along with evaluation results for each, including the binary independence indexing (BII) model, which is a generalized version of the Maron and Kuhns indexing model.
Overview of the TREC-9 Web Track.
177 Citations2000David Hawking
Information Processing & ManagementEngineering a multi-purpose test collection for Web retrieval experiments
174 Citations2003Peter Bailey, Nick Craswell +1 more
It is confirmed that WT10g contains exploitable link information using a site (homepage) finding experiment and the results show that, on this task, Okapi BM25 works better on propagated link anchor text than on full text.
The Computer JournalA Comparison of Text Retrieval Models
138 Citations1992Howard R. Turtle, W. Bruce Croft
This paper introduces a recent form of the probabilistic model based on inference networks, and shows how the vector space and exact-match models can be described in this framework.
Overview of the fifth text REtrieval conference (TREC-5)
134 Citations1996Ellen M. Voorhees, Donna Harman
Presentation des objectifs, des methodes (la routing task, la tâche adhoc, la comparaison des differents systemes, les concepts de confusion, de filtrage, de multilinguisme, de traitement du langage naturel, evaluation de the pertinence), and des resultats issus des traitements manuels and automatiques des differentes equipes presentes a TREC-5.
Relevance feedback and inference networks
122 Citations1993David Haines, W. Bruce Croft
The inference network model introduced by Turtle and Croft is extended to include relevance feedback techniques, and the difference between relevance feedback on text abstracts and full text collections is studied.
Journal of the American Society for Information ScienceA probabilistic approach to automatic keyword indexing. Part II. An algorithm for probabilistic indexing
105 Citations1975Stephen P. Harter
An algorithm defining a measure of indexability is developed-a measure intended to reflect the relative significance of words in documents that is found to consistently produce indexes superior to those produced by another measure which had previously been identified in the literature as producing the best results.
Journal of the ACMFoundations of Probabilistic and Utility-Theoretic Indexing
97 Citations1978William S. Cooper, M. E. Maron
The present paper derives explicit decision rules of both kinds from a common conceptual and mathematical foundation and is a unified theory of indexing.
Term-specific smoothing for the language modeling approach to information retrieval
90 Citations2002Djoerd Hiemstra
The new language modeling approach is shown to explain a number of practical facts of today's information retrieval systems that are not very well explained by the current state of information retrieval theory, including stop words, mandatory terms, coordination level ranking and retrieval using phrases.
Journal of DocumentationINFORMATION RETRIEVAL BY LOGICAL IMAGING
90 Citations1995Fábio Crestani, C. J. van Rijsbergen
This work proposes an approach based on a completely different assumption: ‘a term is a possible world’ which enables the exploitation of term‐term relationships which are estimated using an information theoretic measure.
Interactive Internet search
86 Citations2000Peter Bruza, Robert McArthur +1 more
Search effectiveness when using query-based Internet search, directory-based search and phrase-based query reformulation assisted search is compared by means of a controlled, user-based experimental study.
A new method of weighting query terms for ad-hoc retrieval
78 Citations1996K. L. Kwok
A new method of automatically weighting query terms for ad-hoc retrieval is introduced that works for short queries that is based on the term usage statistics in a collection and no training is required.
The American StatisticianDe Finetti's Theorem on Exchangeable Variables
74 Citations1976David Heath, William D. Sudderth
The Varieties of Information and Scientific Explanation
74 Citations1999Jaakko Hintikka
The concept of information seems to be strangely neglected by epistemologists and philosophers of language, and philosophers’ attention is called to a few possibilities of correcting it.
Text REtrieval ConferenceINQUERY at TREC-5
72 Citations1996James Allan, James P. Callan +5 more
L'equipe de l'universite du Massachusetts a explore trois techniques de traitement de question : le traitement of the question de base, le traduction de la question centrale ( concept cle), l'analyse du contexte local.
On Semantic Information
69 Citations1970Jaakko Hintikka
In the last couple of decades, a logician or a philosopher has run a risk whenever he has put the term “information” into the title of one of his papers because of the expectation that the paper has something to do with that impressive body of results in communication theory.
Journal of DocumentationON RELEVANCE WEIGHT ESTIMATION AND QUERY EXPANSION
65 Citations1986Stephen Robertson
A Bayesian argument is used to suggest modifications to the Robertson/Sparck Jones relevance weighting formula, to accommodate the addition to the query of terms taken from the relevant documents identified during the search.
CWI's Institutional Repository (Centrum Wiskunde & Informatica)Relating the new language models of information retrieval to the traditional retrieval models
58 Citations2000Djoerd Hiemstra, Arjen P. de Vries
ACM Transactions on Information SystemsExperiments with a component theory of probabilistic information retrieval based on single terms as document components
53 Citations1990K. L. Kwok
A component theory of information retrieval using single content terms as component for queries and documents was reviewed and experimented with and performed substantially better than Croft's model because of the highly specific nature of document-focused feedback.
Journal of the ACMComputational Complexity and Probability Constructions
51 Citations1970David G. Willis
Using any universal Tur ing machine as a basis, it is possible to cons t ruc t an infinite number of increas ingly accurate computable probabil i ty measures which are independen t of any p robab i l i ty assumpt ions.
Journal of the Royal Statistical Society Series A (General)The Bivariate Generalized Waring Distribution and its Application to Accident Theory
46 Citations1984Evdokia Xekalaki
Journal of the ACMOperations Research Applied to Document Indexing and Retrieval Decisions
42 Citations1977Abraham Bookstein, Don Kraft
The earher model is extended to include interactions among terms, which allows one to decide whether to retrieve a document by taking into consideration occurrences of all the words in the text.
Text REtrieval ConferenceFUB at TREC-10 Web track: A probabilistic framework for topic relevance term weighting
42 Citations2001Giambattista Amati, Claudio Carpineto +1 more
This approach endeavours to determine the weight of a word within a document in a purely theoretic way as a combination of probability distributions, with the goal of reducing as much as possible the number of parameters which must be learned and tuned from relevance assessments on training test collections.
Information SystemsA probabilistic inference model for information retrieval
38 Citations1991S. K. M. Wong, Yiyu Yao
It is argued that some of the problems presented in the conventional probabilistic models may be resolved if one takes the epistemological view instead of the aleatory view of probability.
Inferring query models by computing information flow
36 Citations2002Peter Bruza, Dawei Song
An alternative, non-probabilistic approach to query modelling whereby the strength of information flow is computed between a query Q and a term w, a reflection of how strongly w is informationally contained within the query Q.
American DocumentationAn experiment in automatic indexing
35 Citations1965Fred J. Damerau
The results of the experiment are encouraging, although not definitive because any index set chosen must be tested by using it for retrieval from a large collection.
N-Poisson document modelling
29 Citations1992Eugene L. Margulis
A practical algorithm for determining if a certain word is distributed acording to an nP distribution and computing the distribution parameters is described and it was found that over 70% of frequently occurring words and terms indeed behave according to the nP distributions.
Journal of Applied ProbabilityPoisson mixtures and quasi-infinite divisibility of distributions
27 Citations1979Prem Puri, Charles M. Goldie
Communication in Statistics- Theory and MethodsParameter estimation for a word frequency distribution based on occupancy theory
25 Citations1986H. S. Sichel
American DocumentationStatistical generation of a technical vocabulary
23 Citations1968Don C. Stone, Morris Rubinoff
The results of an experiment in the use of statistical techniques for extracting a technical vocabulary from document texts are presented and discussed.
Information Processing & ManagementOptimum probability estimation from empirical distributions
21 Citations1989Norbert Fuhr, Hubert Hüther
An optimum estimate for binary features is defined which can be applied to various typical estimation problems in IR and a method for computing this estimate using empirical data is described.
Automatic query expansion based on divergence
19 Citations2001Deng Cai, C. J. van Rijsbergen +1 more
The basic principles and ideas on which the study is based are described, and a theoretical framework is established, which allows the comparison and evaluation of different term scoring functions for identifying good terms for query expansion.
Lecture notes in computer scienceTerm Frequency Normalization via Pareto Distributions
18 Citations2002Giambattista Amati, Cornelis J. van Rijsbergen
Preliminary results show that the unique parameter of the framework can be eliminated in favour of the the term frequency normalization derived by the Paretian law.
Lecture notes in computer sciencePhilosophical issues in Kolmogorov complexity
13 Citations1992Ming Li, Paul Vitányi
This essay is not meant to be an exhaustive survey of the subject, not even of the recent results, but to convey to the reader some appealing philosophical ideas by just picking up some pretty shells deposited on the shore by the sea of applications of Kolmogorov complexity.
Text REtrieval ConferenceTREC-2 routing and ad-hoc retrieval evaluation using the INQUERY system
12 Citations1993W. Bruce Croft, James P. Callan +1 more
The general approach to achieve these goals has been to use improved representations of text and information needs in the framework of a new model of retrieval that uses Bayesian netwoks to describe how text and queries should be uses to identify relevant document.
University of Glasgow at the Web track of TREC 2002
9 Citations2002Vassilis Plachouras, Iadh Ounis +2 more
The aim of the participation in the topic distillation and the named page finding tasks of the Web track is the evaluation of a well-founded modular probabilistic framework for Web Information Retrieval, which integrates content and link analyses.
Studies in fuzziness and soft computingProbabilistic Learning by Uncertainty Sampling with Non-Binary Relevance
3 Citations2000Giambattista Amati, Fábio Crestani
This work presents a learning model for probabilistic learning in information retrieval and information filtering which is based on the concept of “uncertainty sampling” and shows how this new learning model could be evaluated using collections with non-binary relevance assessments.
…
