Statistical language models for information retrieval
Published 1 January 2007Open access
ChengXiang Zhai
Citations219
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
Abstract
Statistical language models have recently been successfully applied to many information retrieval problems. A great deal of recent work has shown that statistical language models not only achieve superior empirical performance, but also facilitate parameter tuning and provide a more principled way, in general, for modeling various kinds of complex and non-traditional retrieval problems.
Keywords
Computer Science
GeneticsInference of Population Structure Using Multilocus Genotype Data
34,260 Citations2000Jonathan K. Pritchard, Matthew Stephens +1 more
A model-based clustering method for using multilocus genotype data to infer population structure and assign individuals to populations that can be applied to most of the commonly used genetic markers, provided that they are not closely linked.
Proceedings of the IEEEA tutorial on hidden Markov models and selected applications in speech recognition
22,785 Citations1989L. R. Rabiner
Journal of the American Society for Information ScienceIndexing by latent semantic analysis
12,677 Citations1990Scott Deerwester, Susan Dumais +3 more
Information Processing & ManagementTerm-weighting approaches in automatic text retrieval
9,532 Citations1988Gerard Salton, Chris Buckley
This paper summarizes the insights gained in automatic term weighting, and provides baseline single term indexing models with which other more elaborate content analysis procedures can be compared.
Communications of the ACMA vector space model for automatic indexing
7,434 Citations1975Gerard Salton, Anita M.-Y. Wong +1 more
An approach based on space density computations is used to choose an optimum indexing vocabulary for a collection of documents, demonstating the usefulness of the model.
IEEE Transactions on Information TheoryDivergence measures based on the Shannon entropy
5,027 Citations1991Jinfeng Lin
A novel class of information-theoretic divergence measures based on the Shannon entropy is introduced, which do not require the condition of absolute continuity to be satisfied by the probability distributions involved and are established in terms of bounds.
ACM Transactions on Information SystemsCumulated gain-based evaluation of IR techniques
4,634 Citations2002Kalervo Järvelin, Jaana Kekäläinen
This article proposes several novel measures that compute the cumulative gain the user obtains by examining the retrieval result up to a given ranked position, and test results indicate that the proposed measures credit IR methods for their ability to retrieve highly relevant documents and allow testing of statistical significance of effectiveness differences.
Journal of DocumentationA STATISTICAL INTERPRETATION OF TERM SPECIFICITY AND ITS APPLICATION IN RETRIEVAL
4,442 Citations1972Karen Spärck Jones
It is argued that terms should be weighted according to collection frequency, so that matches on less frequent, more specific, terms are of greater value than matches on frequent terms.
Optimizing search engines using clickthrough data
3,898 Citations2002Thorsten Joachims
The goal of this paper is to develop a method that utilizes clickthrough data for training, namely the query-log of the search engine in connection with the log of links the users clicked on in the presented ranking.
A comparison of event models for naive bayes text classification
3,224 Citations1998Andrew McCallum, Kamal Nigam
It is found that the multi-variate Bernoulli performs well with small vocabulary sizes, but that the multinomial performs usually performs even better at larger vocabulary sizes--providing on average a 27% reduction in error over the multi -variateBernoulli model at any vocabulary size.
Learning to rank using gradient descent
2,759 Citations2005Chris Burges, Tal Shaked +5 more
RankNet is introduced, an implementation of these ideas using a neural network to model the underlying ranking function, and test results on toy data and on data from a commercial internet search engine are presented.
ACM SIGIR ForumA Language Modeling Approach to Information Retrieval
2,532 Citations2017Jay Ponte, W. Bruce Croft
It will be shown that probabilistic methods can be used to predict topic changes in the context of the task of new event detection and provide further proof of concept for the use of language models for retrieval tasks.
Dynamic topic models
2,322 Citations2006David M. Blei, John Lafferty
A family of probabilistic time series models is developed to analyze the time evolution of topics in large document collections, and dynamic topic models provide a qualitative window into the contents of a large document collection.
ACM SIGIR ForumThe Use of MMR, Diversity-Based Reranking for Reordering Documents and Producing Summaries
2,210 Citations2017Jaime Carbonell, Jade Goldstein
A method for combining query-relevance with information-novelty in the context of text retrieval and summarization and preliminary results indicate some benefits for MMR diversity ranking in document retrieval and in single document summarization.
arXiv (Cornell University)Probabilistic Latent Semantic Analysis
2,092 Citations2013Thomas Hofmann
This work proposes a widely applicable generalization of maximum likelihood model fitting by tempered EM, based on a mixture decomposition derived from a latent class model which results in a more principled approach which has a solid foundation in statistics.
Lecture notes in computer scienceNaive (Bayes) at forty: The independence assumption in information retrieval
2,092 Citations1998David Lewis
The naive Bayes classifier, currently experiencing a renaissance in machine learning, has long been a core technique in information retrieval, and some of the variations used for text retrieval and classification are reviewed.
Journal of the American Society for Information ScienceRelevance weighting of search terms
2,068 Citations1976Stephen Robertson, Karen Spärck Jones
This paper examines statistical techniques for exploiting relevance information to weight search terms using information about the distribution of index terms in documents in general and shows that specific weighted search methods are implied by a general probabilistic theory of retrieval.
A statistical approach to machine translation
1,699 Citations1990Peter F. Brown, John Cocke +6 more
The application of the statistical approach to translation from French to English and preliminary results are described and the results are given.
IEEE Transactions on Acoustics Speech and Signal ProcessingEstimation of probabilities from sparse data for the language model component of a speech recognizer
1,645 Citations1987Slava M. Katz
The model offers, via a nonlinear recursive procedure, a computation and space efficient solution to the problem of estimating probabilities from sparse data, and compares favorably to other proposed methods.
ACM SIGIR ForumA Study of Smoothing Methods for Language Models Applied to Ad Hoc Information Retrieval
1,571 Citations2017ChengXiang Zhai, John Lafferty
This paper examines the sensitivity of retrieval performance to the smoothing parameters and compares several popular smoothing methods on different test collection.
Journal of the American Society for Information ScienceImproving retrieval performance by relevance feedback
1,504 Citations1990Gerard Salton, Chris Buckley
Relevance feedback is an automatic process, introduced over 20 years ago, designed to produce query formulations following an initial retrieval operation to demonstrate the effectiveness of the various methods.
ACM SIGIR ForumRelevance-Based Language Models
1,449 Citations2017Victor Lavrenko, W. Bruce Croft
This work proposes a novel technique for estimating a relevance model with no training data and demonstrates that it can produce highly accurate relevance models, addressing important notions of synonymy and polysemy.
ACM SIGIR ForumQuary Expansion Using Local and Global Document Analysis
1,269 Citations2017Jinxi Xu, W. Bruce Croft
It is shown that, although global analysis haa some advantages, local analysia is generally more effective than global techniques and using global analysis techniques.
ACM Transactions on Information SystemsA study of smoothing methods for language models applied to information retrieval
1,212 Citations2004ChengXiang Zhai, John Lafferty
Evaluation on five different databases and four types of queries indicates that the two-stage smoothing method with the proposed parameter estimation methods consistently gives retrieval performance that is close to or better than the best results achieved using a single smoothing methods and exhaustive parameter search on the test data.
LDA-based document models for ad-hoc retrieval
1,086 Citations2006Xing Wei, W. Bruce Croft
This paper proposes an LDA-based document model within the language modeling framework, and evaluates it on several TREC collections, and shows that improvements over retrieval using cluster-based models can be obtained with reasonable efficiency.
Journal of DocumentationTHE PROBABILITY RANKING PRINCIPLE IN IR
1,071 Citations1977Stephen Robertson
It is shown that the principle that documents should be ranked in order of the probability of relevance or usefulness can be justified under certain assumptions, but that in cases where these assumptions do not hold, the principle is not valid.
Modeling annotated data
1,070 Citations2003David M. Blei, Michael I. Jordan
Three hierarchical probabilistic mixture models which aim to describe annotated data with multiple types, culminating in correspondence latent Dirichlet allocation, a latent variable model that is effective at modeling the joint distribution of both types and the conditional distribution of the annotation given the primary type.
Communications of the ACMExtended Boolean information retrieval
1,038 Citations1983Gerard Salton, Edward A. Fox +1 more
A new, extended Boolean information retrieval system is introduced which is intermediate between the Boolean system of query processing and the vector processing model, and Laboratory tests indicate that the extended system produces better retrieval output than either the Boolean or thevector processing systems.
Distributional clustering of English words
994 Citations1993Fernando Pereira, Naftali Tishby +1 more
Information Processing & ManagementA probabilistic model of information retrieval: development and comparative experiments
947 Citations2000Karen Spärck Jones, Steve Walker +1 more
Neural Information Processing SystemsHierarchical Topic Models and the Nested Chinese Restaurant Process
938 Citations2003Thomas L. Griffiths, Michael I. Jordan +2 more
A Bayesian approach is taken to generate an appropriate prior via a distribution on partitions that allows arbitrarily large branching factors and readily accommodates growing data collections.
Journal of the ACMOn Relevance, Probabilistic Indexing and Information Retrieval
901 Citations1960M. E. Maron, J. L. Kuhns
The paper suggests an interpretation of the whole library problem as one where the request is considered as a clue on the basis of which the library system makes a concatenated statistical inference in order to provide as an output an ordered list of those documents which most probably satisfy the information needs of the user.
ACM Transactions on Information SystemsProbabilistic models of information retrieval based on measuring the divergence from randomness
883 Citations2002Gianni Amati, Cornelis J. van Rijsbergen
A framework for deriving probabilistic models of Information Retrieval using term-weighting models obtained in the language model approach by measuring the divergence of the actual term distribution from that obtained under a random process is introduced.
ACM SIGIR ForumPivoted Document Length Normalization
863 Citations2017Amit Singhal, Chris Buckley +1 more
Pivoted normalization is presented, a technique that can be used to modify any normalization function thereby reducing the gap between the relevance and the retrieval probabilities, and two new normalization functions--pivoted unique normalization and piuotert byte size normalization are presented.
A Markov random field model for term dependencies
849 Citations2005Donald Metzler, W. Bruce Croft
A novel approach is developed to train the model that directly maximizes the mean average precision rather than maximizing the likelihood of the training data, and significant improvements are possible by modeling dependencies, especially on the larger web collections.
Topic sentiment mixture
813 Citations2007Qiaozhu Mei, Xu Ling +3 more
The proposed Topic-Sentiment Mixture (TSM) model can reveal the latent topical facets in a Weblog collection, the subtopics in the results of an ad hoc query, and their associated sentiments and could also provide general sentiment models that are applicable to any ad hoc topics.
Model-based feedback in the language modeling approach to information retrieval
799 Citations2001ChengXiang Zhai, John Lafferty
This paper proposes and evaluates two different approaches to updating a query language model based on feedback documents, one based on a generative probabilistic model of feedback documents and onebased on minimization of the KL-divergence over feedback documents.
Medical Entomology and ZoologySearch Engines: Information Retrieval in Practice
782 Citations2009Bruce Croft, Donald Metzler +1 more
This text provides the background and tools needed to evaluate, compare and modify search engines and numerous programming exercises make extensive use of Galago, a Java-based open source search engine.
ACM SIGIR ForumDocument Language Models, Query Models, and Risk Minimization for Information Retrieval
774 Citations2017John Lafferty, ChengXiang Zhai
A framework for information retrieval that combines document models and query models using a probabilistic ranking function based on Bayesian decision theory is presented and an operational retrieval model that extends recent developments in the language modeling approach to information retrieval is suggested.
Proceedings of the IEEETwo decades of statistical language modeling: where do we go from here?
726 Citations2000Roni Rosenfeld
A Bayesian approach to integration of linguistic theories with data is argued for inStatistical language models estimate the distribution of various natural language phenomena for the purpose of speech recognition and other language technologies.
Distributional clustering of words for text classification
683 Citations1998Lee D. Baker, Andrew Kachites McCallum
This paper describes the application of Distributional Clustering to document classi(cid:12)cation and shows that it can reduce the feature dimensionality by three orders of magnitude and lose only 2% accuracy, better than Latent Semantic In-dexing, class-based clustering, feature selection by mutual information, or Markov-blanket-based feature selection.
ACM SIGIR ForumInformation Retrieval as Statistical Translation
631 Citations2017Adam Berger, John Lafferty
A simple, well motivated model of the document-to-query translation process is proposed, and an algorithm for learning the parameters of this model in an unsupervised manner from a collection of documents is described.
Pachinko allocation
612 Citations2006Wei Li, Andrew McCallum
Improved performance of PAM is shown in document classification, likelihood of held-out data, the ability to support finer-grained topics, and topical keyword coherence.
Probabilistic author-topic models for information discovery
581 Citations2004Mark Steyvers, Padhraic Smyth +2 more
The methodology is applied to a large corpus of 160,000 abstracts and 85,000 authors from the well-known CiteSeer digital library, and a model with 300 topics is learned using a Markov chain Monte Carlo algorithm.
ACM Transactions on Information SystemsEvaluation of an inference network-based retrieval model
577 Citations1991Howard R. Turtle, W. Bruce Croft
Network representations show promise as mechanisms for inferring probable relationships between documents and queries and have been used in information retrieval since at least the early 1960s.
ACM SIGIR ForumBeyond Independent Relevance
567 Citations2015ChengXiang Zhai, William W. Cohen +1 more
A framework for evaluating subtopic retrieval is proposed which generalizes the traditional precision and recall metrics by accounting for intrinsic topic difficulty as well as redundancy in documents and a maximal marginal relevance (MMR) ranking strategy is proposed.
Formal models for expert finding in enterprise corpora
552 Citations2006Krisztian Balog, Leif Azzopardi +1 more
This work presents two general strategies to expert searching given a document collection which are formalized using generative probabilistic models, and shows that the second strategy consistently outperforms the first.
Database and Expert Systems ApplicationsThe INQUERY Retrieval System
526 Citations1992James P. Callan, W. Bruce Croft +1 more
A retrieval system (INQUERY) that is based on a probabilistic retrieval model and provides support for sophisticated indexing and complex query formulation is described.
Computer Speech & LanguageOn structuring probabilistic dependences in stochastic language modelling
516 Citations1994Hermann Ney, U. Essen +1 more
The problem of stochastic language modelling is studied from the viewpoint of introducing suitable structures into the conditional probability distributions, and nonlinear interpolation as an alternative to linear interpolation; equivalence classes for word histories and single words; cache memory and word associations are considered.
A hidden Markov model information retrieval system
469 Citations1999David R. Miller, Tim Leek +1 more
A novel method for performing blind feedback in the HMM framework, a more complex HMM that models bigram production, and several other algorithmic re nements form a state-of-the-art retrieval system that ranked among the best on the TREC-7 ad hoc retrieval task.
Journal of DocumentationA THEORETICAL BASIS FOR THE USE OF CO‐OCCURRENCE DATA IN INFORMATION RETRIEVAL
465 Citations1977C. J. van Rijsbergen
This paper provides a foundation for a practical way of improving the effectiveness of an automatic retrieval system by measuring the extent of the dependence between index terms and using it to construct a non‐linear weighting function.
Expectation-propagation for the generative aspect model
442 Citations2002Thomas P. Minka, John Lafferty
Context-sensitive information retrieval using implicit feedback
437 Citations2005Xuehua Shen, Bin Tan +1 more
This paper proposes several context-sensitive retrieval algorithms based on statistical language models to combine the preceding queries and clicked document summaries with the current query for better ranking of documents.
Journal of DocumentationUSING PROBABILISTIC MODELS OF DOCUMENT RETRIEVAL WITHOUT RELEVANCE INFORMATION
436 Citations1979W. Bruce Croft, David J. Harper
This paper considers the situation where no relevance information is available, that is, at the start of the search, based on a probabilistic model, and proposes strategies for the initial search and an intermediate search.
Journal of the American Society for Information ScienceA theory of term importance in automatic text analysis
388 Citations1975Gerard Salton, Chul‐Su Yang +1 more
Most existing automatic content analysis and indexing techniques are based on word frequency characteristics applied largely in an ad hoc manner, but terms exhibiting high occurence frequencies in individual documents are often useful for high recall performance, whereas terms with low frequency in the whole collection are useful forhigh precision.
Kluwer Academic Publishers eBooksDistributed Information Retrieval
387 Citations2005Jamie Callan
A broad and diverse group of experimental results is presented to demonstrate that the algorithms are effective, efficient, robust, and scalable.
A formal study of information retrieval heuristics
352 Citations2004Hui Fang, Tao Tao +1 more
A formal study of retrieval heuristics is presented and it is found that the empirical performance of a retrieval formula is tightly related to how well it satisfies basic desirable constraints.
Foundations and Trends® in Information RetrievalStatistical Language Models for Information Retrieval A Critical Review
292 Citations2008ChengXiang Zhai
The purpose of this survey is to systematically and critically review the existing work in applying statistical language models to information retrieval, summarize their contributions, and point out outstanding challenges.
Modeling word burstiness using the Dirichlet distribution
277 Citations2005Rasmus Elsborg Madsen, David Kauchak +1 more
The Dirichlet compound multinomial model (DCM) is proposed, which has one additional degree of freedom, which allows it to capture burstiness of words in a document, and performance is comparable to that obtained with multiple heuristic changes to the mult inomial model.
A cross-collection mixture model for comparative text mining
275 Citations2004ChengXiang Zhai, Atulya Velivelli +1 more
A generative probabilistic mixture model is proposed for comparative text mining that simultaneously performs cross-collection clustering and within- collection clustering, and can be applied to an arbitrary set of comparable text collections.
The Importance of Prior Probabilities for Entry Page Search
268 Citations2002Wessel Kraaij, Thijs Westerveld +1 more
Three non-content features of web pages are explored: page length, number of incoming links and URL form, which proved to be a good predictor of entry page search results.
Dependence language model for information retrieval
261 Citations2004Jianfeng Gao, Jian‐Yun Nie +2 more
The linkage of a query is integrated as a hidden variable, which expresses the term dependencies within the query as an acyclic, planar, undirected graph, which extends the basic language modeling approach based on unigram by relaxing the independence assumption.
Combining document representations for known-item search
237 Citations2003Paul Ogilvie, Jamie Callan
This paper investigates the pre-conditions for successful combination of document representations formed from structural markup for the task of known-item search, and presents a mixture-based language model to investigate several hypotheses.
Mining long-term search history to improve search accuracy
233 Citations2006Bin Tan, Xuehua Shen +1 more
This paper study statistical language modeling based methods to mine contextual information from long-term search history and exploit it for a more accurate estimate of the query language model.
Text REtrieval ConferenceExperiments Using the Lemur Toolkit.
223 Citations2001Paul Ogilvie, James P. Callan
This paper describes experiments with Lemur Toolkit, where the author participated in the ad-hoc retrieval task of the Web Track.
Cross-lingual relevance models
217 Citations2002Victor Lavrenko, Martin Choquette +1 more
A formal model of Cross-Language Information Retrieval that does not rely on either query translation or document translation and integrates popular techniques of disambiguation and query expansion in a unified formal framework is proposed.
Regularized estimation of mixture models for robust pseudo-relevance feedback
209 Citations2006Tao Tao, ChengXiang Zhai
A more robust method for pseudo feedback based on statistical language models to integrate the original query with feedback documents in a single probabilistic mixture model and regularize the estimation of the language model parameters in the model so that the information in the feedback documents can be gradually added to the originalquery.
Probabilistic Relevance Models Based on Document and Query Generation
209 Citations2003John Lafferty, ChengXiang Zhai
A unified account of the probabilistic semantics underlying the language modeling approach and the traditional Probabilistic model for information retrieval is given, showing that the two approaches can be viewed as being equivalent probabilistically.
ACM Transactions on Information SystemsA probabilistic learning approach for document indexing
202 Citations1991Norbert Fuhr, Chris Buckley
A method for probabilistic document indexing using relevance feedback data that has been collected from a set of queries based on three new concepts, which allows the integration of new text analysis and knowledge-based methods in this approach as well as the consideration of document structures or different types of terms.
Journal of the American Society for Information ScienceA probabilistic approach to automatic keyword indexing. Part I. On the Distribution of Specialty Words in a Technical Literature
198 Citations1975Stephen P. Harter
A mixture of two Poisson distributions is examined in detail as a model of specialty word distribution and a measure intended to identify specialty words, consistent with the 2-Poisson model, is proposed and evaluated.
Mining correlated bursty topic patterns from coordinated text streams
198 Citations2007Xuanhui Wang, ChengXiang Zhai +2 more
A general probabilistic algorithm which can effectively discover correlated bursty patterns and their bursty periods across text streams even if the streams have completely different vocabularies is proposed, which can be applied to any coordinated text streams to discover correlated topic patterns that burst in multiple streams in the same period.
Lecture notes in computer scienceProbabilistic Models for Expert Finding
193 Citations2007Hui Fang, ChengXiang Zhai
This paper proposes and develops a general probabilistic framework for studying expert finding problem and derive two families of generative models (candidate generation models and topic generation models) from the framework that subsume most existing language models proposed for expert finding.
Two-stage language models for information retrieval
193 Citations2002ChengXiang Zhai, John Lafferty
Evaluation on five different databases and four types of queries indicates that the two-stage smoothing method with the proposed parameter estimation methods consistently gives retrieval performance that is close to---or better than---the best results achieved using a single smoothed method and exhaustive parameter search on the test data.
ACM Transactions on Information SystemsOn modeling information retrieval with probabilistic inference
185 Citations1995S. K. M. Wong, Yiyu Yao
This article examines and extends the logical models of information retrieval in the context of probability theory, and the fundamental notions of term weights and relevance are given probabilistic interpretations.
International Journal on Digital LibrariesA probabilistic justification for using tf×idf term weighting in information retrieval
183 Citations2000Djoerd Hiemstra
The paper shows that the new probabilistic interpretation of tf×idf term weighting might lead to better understanding of statistical ranking mechanisms, for example by explaining how they relate to coordination level ranking.
Query expansion using random walk models
181 Citations2005Kevyn Collins‐Thompson, Jamie Callan
A Markov chain framework that combines multiple sources of knowledge on term associations is described and the effectiveness of the model is evaluated by examining the accuracy and robustness of the expansion methods, and the relative effectiveness of various sources of term evidence is investigated.
PageRank without hyperlinks
179 Citations2005Oren Kurland, Lillian Lee
A number of re-ranking criteria based on measures of centrality in the graphs formed by generation links are studied, and it is shown that integrating centrality into standard language-model-based retrieval is quite effective at improving precision at top ranks.
Corpus structure, language models, and ad hoc information retrieval
171 Citations2004Oren Kurland, Lillian Lee
Integrating word relationships into language models
166 Citations2005Guihong Cao, Jian‐Yun Nie +1 more
The results show that the model achieves substantial and significant improvements with respect to the models without these relationships, and clearly shows the benefit of integrating word relationships into language models for IR.
Journal of DocumentationAN EVALUATION OF FEEDBACK IN DOCUMENT RETRIEVAL USING CO‐OCCURRENCE DATA
160 Citations1978David J. Harper, C. J. van Rijsbergen
This paper reports experiments with a term weighting model incorporating relevance information in which it is assumed that index terms are distributed dependently and argues that if high recall searches are required, relevance feedback based on the modified dependence model may be superior to the widely used Boolean search.
Information Processing & ManagementA risk minimization framework for information retrieval
149 Citations2005ChengXiang Zhai, John Lafferty
A probabilistic information retrieval framework in which the retrieval problem is formally treated as a statistical decision problem, and how this framework can unify existing retrieval models and accommodate systematic development of new retrieval models is discussed.
Document expansion for speech retrieval
146 Citations1999Amit Singhal, Fernando Pereira
Methods of document expansion for a speech retrieval document by a recognizer using a database of vectors of automatic transcriptions of documents is accessed and the vectors are truncated by removing all terms that are not recognizable by the recognizer to create truncated vectors.
A language modeling framework for resource selection and results merging
144 Citations2002Luo Si, Rong Jin +2 more
This paper extends the language modeling approach to integrate resource selection, ad-hoc searching, and merging of results from different text databases into a single probabilistic retrieval model, designed primarily for Intranet environments.
On relevance weights with little relevance information
144 Citations1997Stephen Robertson, Steve Walker
This research highlights the need to understand more fully the role of emotion in the development of Syetema, as well as the role that language and social media have in this process.
A general language model for information retrieval (poster abstract)
143 Citations1999Fei Song, W. Bruce Croft
According to this new paradigm, each document is viewed as a language sample, and a query as a generation process, retrieved documents are ranked based on the probabilities of producing a query from the corresponding language models of these documents.
Twenty-One at TREC-7: ad-hoc and cross-language track
143 Citations1998Djoerd Hiemstra, Wessel Kraaij
This paper describes the official runs of the Twenty-One group for TREC-7 and develops a new weighting algorithm, which outperforms the popular Cornell version of BM25 on the ad-hoc collection and developed a fuzzy matching algorithm to recover from missing translations and spelling variants of proper names.
Language model information retrieval with document expansion
136 Citations2006Tao Tao, Xuanhui Wang +2 more
This paper constructs a probabilistic neighborhood for each document, and expands the document with its neighborhood information, which provides a more accurate estimation of the document model, thus improves retrieval accuracy.
Noun-phrase analysis in unrestricted text for information retrieval
131 Citations1996David A. Evans, ChengXiang Zhai
This paper describes an hybrid approach to the extraction of meaningful subcompounds from complex noun phrases using both corpus statistics and linguistic heuristics and shows that indexing based on such extracted subcompound improves both recall and precision in an information retrieval system.
Parsimonious language models for information retrieval
129 Citations2004Djoerd Hiemstra, Stephen Robertson +1 more
Parsimonious language models explicitly address the relation between levels of language models that are typically used for smoothing, and need fewer (non-zero) parameters to describe the data.
A study of methods for negative relevance feedback
128 Citations2008Xuanhui Wang, Hui Fang +1 more
Experimental results on several TREC collections show that language model based negative feedback methods are generally more effective than those based on vector-space models, and using multiple negative models is an effective heuristic for negative feedback.
Semantic term matching in axiomatic approaches to information retrieval
127 Citations2006Hui Fang, ChengXiang Zhai
This paper shows that semantic term matching can be naturally incorporated into the axiomatic retrieval model through defining the primitive weighting function based on a semantic similarity function of terms, and shows that such extension can be efficiently implemented as query expansion.
A mixture model for contextual text mining
122 Citations2006Qiaozhu Mei, ChengXiang Zhai
A new general probabilistic model for contextual text mining is proposed that can cover several existing models as special cases and can be applied to many interesting mining tasks, such as temporal text mining, spatiotemporal textmining, author-topic analysis, and cross-collection comparative analysis.
Information RetrievalBias and the limits of pooling for large collections
117 Citations2007Chris Buckley, Darrin L. Dimmick +2 more
It is shown that the judgment sets produced by traditional pooling when the pools are too small relative to the total document set size can be biased in that they favor relevant documents that contain topic title words.
The Cluster-Abstraction Model: Unsupervised Learning of Topic Hierarchies from Text Data
117 Citations1999Thomas Hofmann
This paper presents a novel statistical latent class model for text mining and interactive information access, called Cluster-Abstraction Model (CAM), which is purely data driven and utilizes contact-specific word occurrence statistics.
Relevance models for topic detection and tracking
112 Citations2002Victor Lavrenko, James Allan +4 more
It is demonstrated that relevance models result in very substantial improvements over the language modeling baseline, and how the use of relevance modeling makes it possible to choose a single parameter for within- and cross-mode comparisons of stories.
Multi-aspect expertise matching for review assignment
101 Citations2008Maryam Karimzadehgan, ChengXiang Zhai +1 more
This paper studies how to model multiple aspects of expertise and assign reviewers so that they together can cover all subtopics in the document well and proposes three general strategies for solving this problem.
Using query contexts in information retrieval
97 Citations2007Jing Bai, Jian‐Yun Nie +2 more
Both types of context are integrated in an IR model based on language modeling, including context around query and context within query, showing that each of the context factors brings significant improvements in retrieval effectiveness.
Journal of the ACMFoundations of Probabilistic and Utility-Theoretic Indexing
97 Citations1978William S. Cooper, M. E. Maron
The present paper derives explicit decision rules of both kinds from a common conceptual and mathematical foundation and is a unified theory of indexing.
Communications of the ACMApplying Bayesian networks to information retrieval
97 Citations1995Robert Fung, Brendan Del Favero
Information retrieval (IR) is the identification of documents or other units of information in a collection that are relevant to a particular information need.
…
