Naive (Bayes) at forty: The independence assumption in information retrieval
Lecture notes in computer sciencePublished 1 January 1998Open access
David Lewis
Citations2,092
SJR quartileQ2
SJR score0.35
SNIP0.55
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
The naive Bayes classifier, currently experiencing a renaissance in machine learning, has long been a core technique in information retrieval, and some of the variations used for text retrieval and classification are reviewed.
Abstract
The naive Bayes classifier, currently experiencing a renaissance ] in machine learning, has long been a core technique in information retrieval. We review some of the variations of naive Bayes models used for text retrieval and classification, focusing on the distributional assumptions made about word occurrences in documents.
Keywords
Computer Science
Pattern classification and scene analysis
12,643 Citations1973Richard O. Duda, Peter E. Hart
Lecture notes in computer scienceText categorization with Support Vector Machines: Learning with many relevant features
7,925 Citations1998Thorsten Joachims
SVMs achieve substantial improvements over the currently best performing methods and behave robustly over a variety of di-erent learning tasks, eliminating the need for manual parameter tuning.
Machine LearningOn the Optimality of the Simple Bayesian Classifier under Zero-One Loss
3,072 Citations1997Pedro Domingos, Michael J. Pazzani
The Bayesian classifier is shown to be optimal for learning conjunctions and disjunctions, even though they violate the independence assumption, and will often outperform more powerful classifiers for common training set sizes and numbers of attributes, even if its bias is a priori much less appropriate to the domain.
Information Retrieval: Data Structures and Algorithms
2,428 Citations1992William B. Frakes, Ricardo Baeza‐Yates
For programmers and students interested in parsing text, automated indexing, its the first collection in book form of the basic data structures and algorithms that are critical to the storage and retrieval of documents.
Journal of the American Society for Information ScienceRelevance weighting of search terms
2,068 Citations1976Stephen Robertson, Karen Spärck Jones
This paper examines statistical techniques for exploiting relevance information to weight search terms using information about the distribution of index terms in documents in general and shows that specific weighted search methods are implied by a general probabilistic theory of retrieval.
Journal of the American Society for Information ScienceImproving retrieval performance by relevance feedback
1,504 Citations1990Gerard Salton, Chris Buckley
Relevance feedback is an automatic process, introduced over 20 years ago, designed to produce query formulations following an initial retrieval operation to demonstrate the effectiveness of the various methods.
Scaling up the accuracy of Naive-Bayes classifiers: a decision-tree hybrid
1,363 Citations1996Ron Kohavi
A new algorithm, NBTree, is proposed, which induces a hybrid of decision-tree classifiers and Naive-Bayes classifiers: the decision-Tree nodes contain univariate splits as regular decision-trees, but the leaves contain Naïve-Bayesian classifiers.
Journal of the ACMOn Relevance, Probabilistic Indexing and Information Retrieval
901 Citations1960M. E. Maron, J. L. Kuhns
The paper suggests an interpretation of the whole library problem as one where the request is considered as a clue on the basis of which the library system makes a concatenated statistical inference in order to provide as an output an ordered list of those documents which most probably satisfy the information needs of the user.
ACM SIGIR ForumPivoted Document Length Normalization
863 Citations2017Amit Singhal, Chris Buckley +1 more
Pivoted normalization is presented, a technique that can be used to modify any normalization function thereby reducing the gap between the relevance and the retrieval probabilities, and two new normalization functions--pivoted unique normalization and piuotert byte size normalization are presented.
ACM Transactions on Information SystemsEvaluation of an inference network-based retrieval model
577 Citations1991Howard R. Turtle, W. Bruce Croft
Network representations show promise as mechanisms for inferring probable relationships between documents and queries and have been used in information retrieval since at least the early 1960s.
Journal of the ACMAutomatic Indexing: An Experimental Inquiry
552 Citations1961M. E. Maron
The design, execution and evaluation of a modest experimental study aimed at testing empirically one statistical technique for automatic indexed documents according to their subject content are described.
Pivoted document length normalization
545 Citations1996Amit Singhal, Chris Buckley +1 more
Pivoted normalization is presented, a technique that can be used to modify any normalization function thereby reducing the gap between the relevance and the retrieval probabilities and two new normalization functions are presented–-pivoted unique normalization and piuotert byte size nornaahzation.
Computers and the HumanitiesA method for disambiguating word senses in a large corpus
534 Citations1992William A. Gale, Kenneth Church +1 more
The proposed method was designed to disambiguate senses that are usually associated with different topics using a Bayesian argument that has been applied successfully in related tasks such as author identification and information retrieval.
Journal of DocumentationA THEORETICAL BASIS FOR THE USE OF CO‐OCCURRENCE DATA IN INFORMATION RETRIEVAL
465 Citations1977C. J. van Rijsbergen
This paper provides a foundation for a practical way of improving the effectiveness of an automatic retrieval system by measuring the extent of the dependence between index terms and using it to construct a non‐linear weighting function.
Overview of the third text Retrieval conference (TREC-3)
379 Citations1995Donna Harman
Evaluates new technologies in text retrieval with 34 papers: indexing structures, fragmentation schemes, probabilistic retrieval, latent semantic indexing, interactive document retrieval, & much more.
International ACM SIGIR Conference on Research and Development in Information RetrievalProbabilistic models of indexing and searching
313 Citations1980Stephen Robertson, C. J. van Rijsbergen +1 more
There is a considerable body of related work by Salton, Yu and associates on automatic indexing using within-document frequencies of terms.
Evaluating and optimizing autonomous text classification systems
312 Citations1995David Lewis
This work shows how to define what constitutes good effectiveness for binary text classification systems, tune the systems to achieve the highest possible effectiveness, and estimate how the effectiveness changes as new data is processed.
Information Retrieval Systems: Theory and Implementation
289 Citations1997Gerald Kowalski
This chapter discusses information processing systems in detail, including information system evaluation, cataloging and indexing, and the role of text search algorithms in this system.
Communications of the ACMNatural language processing for information retrieval
281 Citations1996David Lewis, Karen Spärck Jones
The paper considers the new opportunities and challenges presented by the user’s ability to search full text directly (rather than e.g. titles and abstracts), and suggests appropriate approaches to doing this, with a focus on the potential role of natural language processing.
Natural Language EngineeringDistribution of content words and phrases in text and language modelling
257 Citations1996Slava M. Katz
The derivation of models describing word distribution in text is based on a linguistic interpretation of the process of text formation, with the probabilities of word occurrence being functions of observable and linguistically meaningful text characteristics.
Context-sensitive learning methods for text categorization
253 Citations1996William W. Cohen, Yoram Singer
Journal of the American Society for Information ScienceA probabilistic approach to automatic keyword indexing. Part I. On the Distribution of Specialty Words in a Technical Literature
198 Citations1975Stephen P. Harter
A mixture of two Poisson distributions is examined in detail as a model of specialty word distribution and a measure intended to identify specialty words, consistent with the 2-Poisson model, is proposed and evaluated.
Springer series in statisticsApplied Bayesian and Classical Inference
183 Citations1984Frederick Mosteller, David L. Wallace
Find loads of the applied bayesian and classical inference book catalogues in this site as the choice of you visiting this page.
Information Processing & ManagementModels for retrieval with probabilistic indexing
177 Citations1989Norbert Fuhr
Three retrieval models for probabilistic indexing are described along with evaluation results for each, including the binary independence indexing (BII) model, which is a generalized version of the Maron and Kuhns indexing model.
Journal of DocumentationAN EVALUATION OF FEEDBACK IN DOCUMENT RETRIEVAL USING CO‐OCCURRENCE DATA
160 Citations1978David J. Harper, C. J. van Rijsbergen
This paper reports experiments with a term weighting model incorporating relevance information in which it is assumed that index terms are distributed dependently and argues that if high recall searches are required, relevance feedback based on the modified dependence model may be superior to the widely used Boolean search.
Using Taxonomy, Discriminants, and Signatures for Navigating in Text Databases
130 Citations1997Soumen Chakrabarti, Byron Dom +2 more
This work uses techniques from statistical pattern recognition to efficiently separate the feature words or discriminants from the noise words at each node of the taxonomy, and builds a multi-level classifier that has a small model size and is very fast.
Journal of the American Society for Information ScienceA probabilistic approach to automatic keyword indexing. Part II. An algorithm for probabilistic indexing
105 Citations1975Stephen P. Harter
An algorithm defining a measure of indexability is developed-a measure intended to reflect the relative significance of words in documents that is found to consistently produce indexes superior to those produced by another measure which had previously been identified in the literature as producing the best results.
Journal of the American Society for Information ScienceA decision theoretic foundation for indexing
91 Citations1975Abraham Bookstein, Don R. Swanson
Though the main purpose of this paper is to provide insights into a very complex process, formulae are developed that may prove to be of value for an automated operating system.
Psychology Press eBooksText-based intelligent Systems
87 Citations2014Paul S. Jacobs
This chapter discusses Text Representation for Intelligent Text Retrieval: A Classification-Oriented View, and Intelligent High-Volume Text Processing Using Shallow, Domain-Specific Techniques.
ACM Transactions on Information SystemsSome inconsistencies and misidentified modeling assumptions in probabilistic information retrieval
73 Citations1995William S. Cooper
Research in the probabilistic theory of information retrieval involves the construction of mathematical models based on statistical assumptions, including the so-called Binary Independence model, which has been seriously misapprehended.
Journal of DocumentationSEARCH TERM RELEVANCE WEIGHTING GIVEN LITTLE RELEVANCE INFORMATION
72 Citations1979Karen Spärck Jones
The tests simulated iterative searching, as in an on‐line system, and show that even very little relevance information can be of considerable value in relation to relevance weighting.
Journal of the American Society for Information ScienceParameter estimation for probabilistic document-retrieval models
54 Citations1988Robert M. Losee
A proposal that parameters of distributions describing the distribution of features in nonrelevant documents be estimated from the parameters of the corresponding distributions of the entire database is tested; the confidence parameter of such an estimate resulting in the highest average precision is given.
Journal of the ACMOperations Research Applied to Document Indexing and Retrieval Decisions
42 Citations1977Abraham Bookstein, Don Kraft
The earher model is extended to include interactions among terms, which allows one to decide whether to retrieve a document by taking into consideration occurrences of all the words in the text.
One term or two?
36 Citations1995Kenneth Church
Document classification by machine
33 Citations1994Louise Guthrie, Elbert A. Walker +1 more
A mathematical model of classification schemes and the one scheme which can be proved optimal among all those based on word frequencies is described and an experiment illustrates the efficacy of this classification method.
Document classification using a finite mixture model
32 Citations1997Hang Li, Kenji Yamanishi
This work treats the problem of classifying documents as that of conducting statistical hypothesis testing over finite mixture models, and employs the EM algorithm to efficiently estimate parameters in a finite mixture model.
Information Processing & ManagementModelling documents with multiple poisson distributions
20 Citations1993Eugene L. Margulis
It was found that over 70% of frequently occurring terms indeed behave according to the nP distributions and the results indicate that the proportion of nP terms is even higher for the collections in which documents have similar length.
Two learning schemes in information retrieval
9 Citations1988C. Yu, Hidenobu Mizuno
Two methods are given to improve weighting schemes by using relevance information of a set of queries to estimate parameter values of two independence models in information retrieval — the binary independence model and the non-binary independence model.
Text REtrieval ConferenceBayesian Inference with Node Aggregation for Information Retrieval.
5 Citations1993Brendan Del Favero, Robert Fung
Research is directed at developing a probabilistic information retrieval architecture that is oriented towards assisting users that have stable information needs in routing large amounts of time-sensitive material and requires modest computational resources.
scholarworks - UTEP (The University of Texas at El Paso)Document Classification by Machine: Theory and Practice
4 Citations1994Guthrie, Louise, Walker, Elbert +1 more
arXiv (Cornell University)Document Classification Using a Finite Mixture Model
4 Citations1997Hang Li, Kenji Yamanishi
