A Survey of Text Classification Algorithms
Published 1 January 2012
Charų C. Aggarwal, ChengXiang Zhai
Citations834
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
A survey of a wide variety of text classification algorithms for a number of diverse domains, including target marketing, medical diagnosis, news group filtering, and document organization is provided.
Abstract
The problem of classification has been widely studied in the data mining, machine learning, database, and information retrieval communities with applications in a number of diverse domains, such as target marketing, medical diagnosis, news group filtering, and document organization. In this paper we will provide a survey of a wide variety of text classification algorithms.
Keywords
Computer Science
Journal of the Royal Statistical Society Series B (Statistical Methodology)Maximum Likelihood from Incomplete Data Via the <i>EM</i> Algorithm
49,657 Citations1977A. P. Dempster, N. M. Laird +1 more
The Nature of Statistical Learning Theory
39,279 Citations1995Vladimir Vapnik
Elements of Information Theory
37,533 Citations2001Thomas M. Cover, Joy A. Thomas
Machine LearningSupport-Vector Networks
33,035 Citations1995Corinna Cortes, Vladimir Vapnik
High generalization ability of support-vector networks utilizing polynomial input transformations is demonstrated and the performance of the support- vector network is compared to various classical learning algorithms that all took part in a benchmark study of Optical Character Recognition.
BiometricsClassification and Regression Trees.
23,841 Citations1984Alexander Gordon, Leo Breiman +3 more
Journal of Computer and System SciencesA Decision-Theoretic Generalization of On-Line Learning and an Application to Boosting
20,320 Citations1997Yoav Freund, Robert E. Schapire
Machine LearningInduction of Decision Trees
14,815 Citations1986J. R. Quinlan
This paper summarizes an approach to synthesizing decision trees that has been used in a variety of systems, and it describes one such system, ID3, in detail, which is described in detail.
Annals of EugenicsTHE USE OF MULTIPLE MEASUREMENTS IN TAXONOMIC PROBLEMS
14,727 Citations1936Ronald Aylmer Fisher
Journal of the American Society for Information ScienceIndexing by latent semantic analysis
12,677 Citations1990Scott Deerwester, Susan Dumais +3 more
Lecture notes in computer scienceText categorization with Support Vector Machines: Learning with many relevant features
7,925 Citations1998Thorsten Joachims
SVMs achieve substantial improvements over the currently best performing methods and behave robustly over a variety of di-erent learning tasks, eliminating the need for manual parameter tuning.
ACM Computing SurveysMachine learning in automated text categorization
7,899 Citations2002Fabrizio Sebastiani
This survey discusses the main approaches to text categorization that fall within the machine learning paradigm and discusses in detail issues pertaining to three different problems, namely, document representation, classifier construction, and classifier evaluation.
Combining labeled and unlabeled data with co-training
5,604 Citations1998Avrim Blum, Tom M. Mitchell
A Comparative Study on Feature Selection in Text Categorization
4,766 Citations1997Yiming Yang, Jan Pedersen
DF thresholding, the simplest method with the lowest cost in computation, can be reliably used instead of IG or CHI when the computation of these measures are too expensive, and strong correlations between the DF, IG and CHI values of a term are found.
Probabilistic latent semantic indexing
3,916 Citations1999Thomas Hofmann
Probabilistic Latent Semantic Indexing is a novel approach to automated document indexing which is based on a statistical latent class model for factor analysis of count data.
A comparison of event models for naive bayes text classification
3,224 Citations1998Andrew McCallum, Kamal Nigam
It is found that the multi-variate Bernoulli performs well with small vocabulary sizes, but that the multinomial performs usually performs even better at larger vocabulary sizes--providing on average a 27% reduction in error over the multi -variateBernoulli model at any vocabulary size.
Machine LearningOn the Optimality of the Simple Bayesian Classifier under Zero-One Loss
3,072 Citations1997Pedro Domingos, Michael J. Pazzani
The Bayesian classifier is shown to be optimal for learning conjunctions and disjunctions, even though they violate the independence assumption, and will often outperform more powerful classifiers for common training set sizes and numbers of attributes, even if its bias is a priori much less appropriate to the domain.
Transductive Inference for Text Classification using Support Vector Machines
2,717 Citations1999Thorsten Joachims
An analysis of why TSVMs are well suited for text classi(cid:12)cation is presented, and an algorithm for training TSVMs e(cid:14)-ciently, handling 10,000 examples and more is proposed.
A re-examination of text categorization methods
2,660 Citations1999Yiming Yang, Xin Liu
The results show that SVM, kNN and LLSF signi cantly outperform NNet and NB when the number of positive training instances per category are small, and that all the methods perform comparably when the categories are over 300 instances.
Integrating classification and association rule mining
2,223 Citations1998Bing Liu, Wynne Hsu +1 more
The integration is done by focusing on mining a special subset of association rules, called class association rules (CARs), and shows that the classifier built this way is more accurate than that produced by the state-of-the-art classification system C4.5.
Lecture notes in computer scienceA desicion-theoretic generalization of on-line learning and an application to boosting
2,213 Citations1995Yoav Freund, Robert E. Schapire
The model studied can be interpreted as a broad, abstract extension of the well-studied on-line prediction model to a general decision-theoretic setting, and it is shown that the multiplicative weight-update Littlestone?Warmuth rule can be adapted to this model, yielding bounds that are slightly weaker in some cases, but applicable to a considerably more general class of learning problems.
Machine LearningBoosTexter: A Boosting-based System for Text Categorization
2,194 Citations2000Robert E. Schapire, Yoram Singer
This work describes in detail an implementation, called BoosTexter, of the new boosting algorithms for text categorization tasks, and presents results comparing the performance of Boos Texter and a number of other text-categorization algorithms on a variety of tasks.
Lecture notes in computer scienceNaive (Bayes) at forty: The independence assumption in information retrieval
2,092 Citations1998David Lewis
The naive Bayes classifier, currently experiencing a renaissance in machine learning, has long been a core technique in information retrieval, and some of the variations used for text retrieval and classification are reviewed.
Journal of the American Society for Information ScienceRelevance weighting of search terms
2,068 Citations1976Stephen Robertson, Karen Spärck Jones
This paper examines statistical techniques for exploiting relevance information to weight search terms using information about the distribution of index terms in documents in general and shows that specific weighted search methods are implied by a general probabilistic theory of retrieval.
Elsevier eBooksNewsWeeder: Learning to Filter Netnews
2,022 Citations1995Ken Lang
The results show that a learning algorithm based on the Minimum Description Length (MDL) principle was able to raise the percentage of interesting articles to be shown to users from 14% to 52% on average.
On Discriminative vs. Generative Classifiers: A comparison of logistic regression and naive Bayes
1,887 Citations2001Andrew Y. Ng, Michael I. Jordan
It is shown, contrary to a widely-held belief that discriminative classifiers are almost always to be preferred, that there can often be two distinct regimes of performance as the training set size is increased, one in which each algorithm does better.
The foundations of cost-sensitive learning
1,835 Citations2001Charles Elkan
It is argued that changing the balance of negative and positive training examples has little effect on the classifiers produced by standard Bayesian and decision tree learning methods, and the recommended way of applying one of these methods is to learn a classifier from the training set and then to compute optimal decisions explicitly using the probability estimates given by the classifier.
Inductive learning algorithms and representations for text categorization
1,465 Citations1998Susan Dumais, John Platt +2 more
A comparison of the effectiveness of five different automatic learning algorithms for text categorization in terms of learning speed, realtime classification speed, and classification accuracy is compared.
IEEE Transactions on Neural NetworksSupport vector machines for spam categorization
1,457 Citations1999Harris Drucker, Donghui Wu +1 more
The use of support vector machines in classifying e-mail as spam or nonspam is studied by comparing it to three other classification algorithms: Ripper, Rocchio, and boosting decision trees, which found SVM's performed best when using binary features.
Machine LearningLearning Quickly When Irrelevant Attributes Abound: A New Linear-Threshold Algorithm
1,376 Citations1988Nick Littlestone
This work presents one such algorithm that learns disjunctive Boolean functions, along with variants for learning other classes of Boolean functions.
A Survey of Opinion Mining and Sentiment Analysis
1,331 Citations2012Bing Liu, Lei Zhang
Sentiment analysis or opinion mining is the computational study of people’s opinions, appraisals, attitudes, and emotions toward entities, individuals, issues, events, topics and their attributes.
MetaCost
1,294 Citations1999Pedro Domingos
A principled method for making an arbitrary classifier cost-sensitive by wrapping a cost-minimizing procedure around it is proposed, called MetaCost, which treats the underlying classifier as a black box, requiring no knowledge of its functioning or change to it.
A Probabilistic Analysis of the Rocchio Algorithm with TFIDF for Text Categorization
1,264 Citations1997Thorsten Joachims
A Probabilistic analysis of the Rocchio relevance feedback algorithm, one of the most popular learning methods from information retrieval, is presented in a text categorization framework and suggests that the probabilistic algorithms are preferable to the heuristic Rocchio classifier.
A Sequential Algorithm for Training Text Classifiers
1,158 Citations1994David Lewis, William A. Gale
An algorithm for sequential sampling during machine learning of statistical classifiers was developed and tested on a newswire text categorization task and reduced by as much as 500-fold the amount of training data that would have to be manually classified to achieve a given level of effectiveness.
A Bayesian Approach to Filtering Junk E-Mail
1,153 Citations1998Mehran Sahami, Susan Dumais +2 more
This work examines methods for the automated construction of filters to eliminate such unwanted messages from a user’s mail stream, and shows the efficacy of such filters in a real world usage scenario, arguing that this technology is mature enough for deployment.
Elsevier eBooksHeterogeneous Uncertainty Sampling for Supervised Learning
1,142 Citations1994David Lewis, Jason Catlett
This work test the use of one classifier (a highly efficient probabilistic one) to select examples for training another (the C4.5 rule induction program) and finds that the uncertainty samples yielded classifiers with lower error rates than random samples ten times larger.
Communications of the ACMExtended Boolean information retrieval
1,038 Citations1983Gerard Salton, Edward A. Fox +1 more
A new, extended Boolean information retrieval system is introduced which is intermediate between the Boolean system of query processing and the vector processing model, and Laboratory tests indicate that the extended system produces better retrieval output than either the Boolean or thevector processing systems.
ACM Transactions on Information SystemsAutomated learning of decision rules for text categorization
867 Citations1994Chidanand Apté, Fred J. Damerau +1 more
It is shown that machine-generated decision rules appear comparable to human performance, while using the identical rule-based representation, and compared with other machine-learning techniques.
Hierarchically Classifying Documents Using Very Few Words
840 Citations1997Daphne Koller, Mehran Sahami
This work proposes an approach that utilizes the hierarchical topic structure to decompose the classification task into a set of simpler problems, one at each node in the classification tree, which can be solved accurately by focusing only on a very small set of features, those relevant to the task at hand.
Hierarchical classification of Web content
805 Citations2000Susan Dumais, Hao Chen
This paper explores the use of hierarchical structure for classifying a large, heterogeneous collection of web content using support vector machine (SVM) classifiers, which have been shown to be efficient and effective for classification, but not previously explored in the context of hierarchical classification.
Semi-supervised Clustering by Seeding
804 Citations2002Sugato Basu, Arindam Banerjee +1 more
Enhanced hypertext categorization using hyperlinks
775 Citations1998Soumen Chakrabarti, Byron Dom +1 more
This work has developed a text classifier that misclassified only 13% of the documents in the well-known Reuters benchmark; this was comparable to the best results ever obtained and its technique also adapts gracefully to the fraction of neighboring documents having known topics.
Distributional clustering of words for text classification
683 Citations1998Lee D. Baker, Andrew Kachites McCallum
This paper describes the application of Distributional Clustering to document classi(cid:12)cation and shows that it can reduce the feature dimensionality by three orders of magnitude and lose only 2% accuracy, better than Latent Semantic In-dexing, class-based clustering, feature selection by mutual information, or Markov-blanket-based feature selection.
Learning to extract symbolic knowledge from the World Wide Web
675 Citations1998Mark Craven, Dan DiPasquo +5 more
The goal of the research described here is to automatically create a computer understandable world wide knowledge base whose content mirrors that of the World Wide Web, and several machine learning algorithms for this task are described.
An evaluation of phrasal and clustered representations on a text categorization task
547 Citations1992David Lewis
It is shown that optimal effectiveness occurs when using only a small proportion of the indexing terms available, and that effectiveness peaks at a higher feature set size and lower effectiveness level for a syntactic phrase indexing than for word-based indexing.
arXiv (Cornell University)An evaluation of Naive Bayesian anti-spam filtering
527 Citations2000Ion Androutsopoulos, John Koutsias +3 more
It is reached that additional safety nets are needed for the Naive Bayesian anti-spam filter to be viable in practice.
Technische Universität Dortmund Eldorado (Technische Universität Dortmund)Text categorization with support vector machines
507 Citations1997Thorsten Joachims
Opinion Mining and Sentiment Analysis
480 Citations2011Bing Liu
This chapter focuses on mining opinions which indicate positive or negative sentiments, which are of great importance for businesses and consumers wanting to find public or consumer opinions on their products and services.
Improving Text Classification by Shrinkage in a Hierarchy of Classes
478 Citations2022Andrew McCallum, Roni Rosenfeld +2 more
This paper shows that the accuracy of a naive Bayes text classi(cid:12)er can be significantly improved by taking advantage of a hierarchy of classes, and adopts an established statistical technique called shrinkage that smoothes parameter estimates of a data-sparse child with its parent in order to obtain more robust parameter estimates.
A comparison of classifiers and document representations for the routing problem
451 Citations1995Hinrich Schütze, David A. Hull +1 more
This paper considers three classification techniques which have decision rules that are derived via explicit error minimization linear discriminant analysis, logistic regression, and neuraf networks, and finds that features based on latent semantic indexing are more effective for techniques such aslinear discriminant anaf-ysis and logistic regressors, which have no way to protect against overfitting.
Journal of Machine Learning ResearchClassification in Networked Data: A Toolkit and a Univariate Case Study
447 Citations2007Sofus A. Macskassy, Foster Provost
The results demonstrate that very simple network-classification models perform quite well---well enough that they should be used regularly as baseline classifiers for studies of learning with networked data.
Feature selection, perception learning, and a usability case study for text categorization
446 Citations1997Hwee Tou Ng, Wei Boon Goh +1 more
An automated learning approach to text categorization based on perception learning and a new feature selection metric, called correlation coefficient, is described and empirical results indicate that this approach outperforms the best published results on this % uters collection.
Margin based feature selection - theory and algorithms
416 Citations2004Ran Gilad-Bachrach, Amir Navot +1 more
This paper introduces a margin based feature selection criterion and applies it to measure the quality of sets of features and devise novel selection algorithms for multi-class classification problems and provide theoretical generalization bound.
Combining classifiers in text categorization
414 Citations1996Leah S. Larkey, W. Bruce Croft
Lecture notes in computer scienceCentroid-Based Document Classification: Analysis and Experimental Results
406 Citations2000Eui-Hong Han, George Karypis
The authors' experiments show that this centroidbased classifier consistently and substantially outperforms other algorithms such as Naive Bayesian, k-nearest-neighbors, and C4.5, on a wide range of datasets.
ACM Transactions on Information SystemsAn example-based mapping method for text categorization and retrieval
405 Citations1994Yiming Yang, Christopher G. Chute
It is evident that the LLSF approach uses the relevance information effectively within human decisions of categorization and retrieval, and achieves a semantic mapping of free texts to their representations in an indexing language.
Learning Rules that Classify E-Mail
400 Citations1996William W. Cohen
Two methods for learning text classifiers are compared on classification problems that might arise in filtering and filing personM e-mail messages: a "traxiitionM IR" method based on TF-IDF weighting, and a new method for learning sets of "keyword-spotting rules" based on the RIPPER rule learning algorithm.
IEEE Transactions on Pattern Analysis and Machine IntelligenceGeneralizing discriminant analysis using the generalized singular value decomposition
363 Citations2004Peg Howland, H. Park
This work examines a number of optimization criteria, and extends their applicability by using the generalized singular value decomposition to circumvent the nonsingularity requirement.
Learning limited dependence Bayesian classifiers
363 Citations1996Mehran Sahami
A framework for characterizing Bayesian classification methods is presented and a general induction algorithm is presented that allows for traversal of this spectrum depending on the available computational power for carrying out induction and its application in a number of domains with different properties.
Node Classification in Social Networks
359 Citations2011Smriti Bhagat, Graham Cormode +1 more
When dealing with large graphs, such as those that arise in the context of online social networks, a subset of nodes may be labeled to indicate demographic values, interest, beliefs or other characteristics of the nodes (users).
ACM Transactions on Information SystemsContext-sensitive learning methods for text categorization
359 Citations1999William W. Cohen, Yoram Singer
RIPPER and sleeping-experts perform extremely well across a wide variety of categorization problems, generally outperforming previously applied learning methods and are viewed as a confirmation of the usefulness of classifiers that represent contextual information.
A study of thresholding strategies for text categorization
354 Citations2001Yiming Yang
Experimental results show that the choice of thresholding strategy can significantly influence the performance of kNN, and that the ``optimal'' strategy may vary by application.
Lecture notes in computer scienceText Categorization Using Weight Adjusted k-Nearest Neighbor Classification
353 Citations2001Eui-Hong Han, George Karypis +1 more
A Weight Adjusted k-Nearest Neighbor (WAKNN) classification that learns feature weights based on a greedy hill climbing technique and two performance optimizations of WAKNN that improve the computational performance by a few orders of magnitude, but do not compromise on the classification quality.
Learning to classify text from labeled and unlabeled documents
330 Citations1998Kamal Nigam, Andrew McCallum +2 more
It is shown that the accuracy of text classifiers trained with a small number of labeled documents can be improved by augmenting this small training set with a large pool of unlabeled documents, and an algorithm is introduced based on the combination of Expectation-Maximization with a naive Bayes classifier.
arXiv (Cornell University)A Bayesian Approach to Learning Bayesian Networks with Local Structure
326 Citations2013David Maxwell Chickering, David Heckerman +1 more
Boosting and Rocchio applied to text filtering
299 Citations1998Robert E. Schapire, Yoram Singer +1 more
This paper discusses two learning algorithms for text filtering: modified Rocchio and a boosting algorithm called AdaBoost, and shows how both algorithms can be adapted to maximize any general utility matrix that associates cost for each pair of machine prediction and correct label.
Combining content and link for classification using matrix factorization
263 Citations2007Shenghuo Zhu, Kai Yu +2 more
This paper aims to design an algorithm that exploits both the content and linkage information, by carrying out a joint factorization on both the linkage adjacency matrix and the document-term matrix, and derives a new representation for web pages in a low-dimensional factor space, without explicitly separating them as content, hub or authority factors.
Decision Support SystemsPartitioning-based clustering for Web document categorization
261 Citations1999Daniel Boley, Maria Gini +7 more
Two new clustering algorithms are introduced that can effectively cluster documents, even in the presence of a very high dimensional feature space, and do not require pre-specified ad hoc distance functions and are capable of automatically discovering document similarities or associations.
Why collective inference improves relational classification
256 Citations2004David Jensen, Jennifer Neville +1 more
This work describes the necessary and sufficient conditions for reduced classification error based on experiments with real and simulated data, and characterizes different types of statistical models used for making inference in relational data.
Unsupervised document classification using sequential information maximization
252 Citations2002Noam Slonim, Nir Friedman +1 more
A novel sequential clustering algorithm which is motivated by the Information Bottleneck method is presented, and it is found to be consistently superior to all the other clustering methods examined, typically by a significant margin.
Using and combining predictors that specialize
239 Citations1997Yoav Freund, Robert E. Schapire +2 more
It is shown how to transform algorithms that assume that all experts are always awake to algorithms that do not require this assumption, and how to derive corresponding loss bounds.
A Statistical Learning Model of Text Classification for Support Vector Machines
238 Citations2001Thorsten Joachims
This model explains why and when SVMs perform well for text classification and connects the statistical properties of text-classification tasks with the generalization performance of a SVM in a quantitative way.
Large scale semi-supervised linear SVMs
233 Citations2006Vikas Sindhwani, S. Sathiya Keerthi
An implementation of Transductive SVM (TSVM) that is significantly more efficient and scalable than currently used dual techniques, for linear classification problems involving large, sparse datasets, and a variant of TSVM that involves multiple switching of labels.
A statistical learning learning model of text classification for support vector machines
223 Citations2001Thorsten Joachims
IEEE Intelligent Systems and their ApplicationsMaximizing text-mining performance
211 Citations1999Sabine Weiß, Chid Apte +5 more
Pattern Recognition LettersOn the exponential value of labeled samples
201 Citations1995Vittorio Castelli, Thomas M. Cover
The first labeled sample reduces the risk from 1 2 to 2R ∗ (1−R∗ ) and subsequent labeled samples in the training set reduce the probability of error exponentially fast to the Bayes risk.
Noise reduction in a statistical approach to text categorization
196 Citations1995Yiming Yang
Noise reduction strategies are proposed and evaluated, including an aggressive removal of “non-informative words” from texts before training; the use of a truncated singular value decomposition to cut off noisy “latent semantic structures” during training.
Learning quickly when irrelevant attributes abound: A new linear-threshold algorithm
193 Citations1987Nick Littlestone
Using a generalized instance set for automatic text categorization
188 Citations1998Wai Lam, Chao Yang Ho
This work proposes a new technique known as the generalized instance set (GIS) algorithm by unifying the strengths of k-NN and linear classifiers and adapting to characteristics of text categorization problems.
Feature selection using linear classifier weights
183 Citations2004Dunja Mladenić, Janez Brank +2 more
Experiments show that feature selection using weights from linear SVMs yields better classification performance than other feature weighting methods when combined with the three explored learning algorithms.
Machine LearningRandom classification noise defeats all convex potential boosters
180 Citations2009Philip M. Long, Rocco A. Servedio
This paper shows that for a broad class of convex potential functions, any such boosting algorithm is highly susceptible to random classification noise, and there is a simple data set of examples which is efficiently learnable by such a booster if there is no noise, but which cannot be learned to accuracy better than 1/2 if there are random classification noises.
SIAM Journal on Matrix Analysis and ApplicationsStructure Preserving Dimension Reduction for Clustered Text Data Based on the Generalized Singular Value Decomposition
174 Citations2003Peg Howland, Moongu Jeon +1 more
This work adapt and extend the discriminant analysis projection used in pattern recognition and shows that by using the generalized singular value decomposition (GSVD), it can achieve the same goal regardless of the relative dimensions of the term-document matrix.
On the collective classification of email "speech acts"
161 Citations2005Vitor R. Carvalho, William W. Cohen
A new text-classification algorithm based on a dependency-network based collective classification method, in which the local classifiers are maximum entropy models based on words and certain relational features, which appears to be consistent across many email acts suggested by prior speech-act theory.
Information RetrievalExploiting Hierarchy in Text Categorization
159 Citations1999Andreas S. Weigend, Erik D. Wiener +1 more
The structure that is present in the semantic space of topics is used in order to improve performance in text categorization: according to their meaning, topics can be grouped together into “meta-topics”, e.g., gold, silver, and copper are all metals.
The Power of Word Clusters for Text Classification
158 Citations2006Noam Slonim, Naftali Tishby
This work applies the information bottleneck method to find word-clusters that preserve the information about document categories and use these clusters as features for classification, and shows that when the training sample is small word clusters can yield significant improvement in classification accuracy.
Information Processing & ManagementThreading electronic mail: A preliminary study
153 Citations1997David Lewis, Kimberly A. Knowles
It is proposed that threading of electronic messages be treated as a language processing task, and that a significant level of threading effectiveness can be achieved by applying standard text matching methods from information retrieval to the textual portions of messages.
Feature reduction for neural network based text categorization
151 Citations2003S.L.Y. Lam, Dik Lun Lee
The proposed and compared four dimensionality reduction techniques to reduce the feature space into an input space of much lower dimension for the neural network classifier showed that the proposed model was able to achieve high categorization effectiveness as measured by precision and recall.
The Role of Unlabeled Data in Supervised Learning
150 Citations2004Tom M. Mitchell
It is argued that models of human and animal learning should consider more strongly the potential role of unlabeled data, and that many natural learning problems fit the problem class identified in this paper.
Graph-based text classification
149 Citations2006Ralitsa Angelova, Gerhard Weikum
A practical hypertext catergorization method using links and incrementally available class information
148 Citations2000Hyo-Jung Oh, Sung Hyon Myaeng +1 more
This paper proposes a practical method for enhancing both the speed and the quality of hypertext categorization using hyperlinks, and achieves up to 18.5% of improvement in effectiveness while reducing the processing time dramatically.
arXiv (Cornell University)Mistake-Driven Learning in Text Categorization
144 Citations1997Ido Dagan, Yael Karov +1 more
This work studies three mistake-driven learning algorithms for a typical task of this nature -- text categorization and presents an algorithm, a variation of Littlestone's Winnow, which performs significantly better than any other algorithm tested on this task using a similar feature set.
The VLDB JournalFast and accurate text classification via multiple linear discriminant projections
140 Citations2003Soumen Chakrabarti, Shourya Roy +1 more
SIMPL is presented, a nearly linear-time classification algorithm that mimics the strengths of SVMs while avoiding the training bottleneck and not only approaches and sometimes exceeds SVM accuracy, but also beats the running time of a popular SVM implementation by orders of magnitude.
ACM SIGIR ForumFeature selection, perceptron learning, and a usability case study for text categorization
140 Citations1997Hwee Tou Ng, Wei Boon Goh +1 more
Deep classification in large-scale text hierarchies
135 Citations2008Gui-Rong Xue, Dikan Xing +2 more
A novel deep-classification approach to categorize Web documents into categories in a large-scale taxonomy using a statistical-language-model based classifier using n-gram features and the structure of the taxonomy is utilized in this stage to improve the performance of classification.
Using Taxonomy, Discriminants, and Signatures for Navigating in Text Databases
130 Citations1997Soumen Chakrabarti, Byron Dom +2 more
This work uses techniques from statistical pattern recognition to efficiently separate the feature words or discriminants from the noise words at each node of the taxonomy, and builds a multi-level classifier that has a small model size and is very fast.
Machine LearningRelational Learning with Statistical Predicate Invention: Better Models for Hypertext
126 Citations2001Mark Craven, Seán Slattery
This work presents a new approach to learning hypertext classifiers that combines a statistical text-learning method with a relational rule learner and demonstrates that this new approach is able to learn more accurate classifiers than either of its constituent methods alone.
…
