Maximizing text-mining performance
IEEE Intelligent Systems and their ApplicationsPublished 1 July 1999
Sabine Weiß, Chid Apte, Fred J. Damerau, David E. Johnson, Frank J. Oles, Thomas Goetz
Citations211
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
Abstract
The authors' adaptive resampling approach surpasses previous decision-tree performance and validates the effectiveness of small, pooled local dictionaries. They demonstrate their approach using the Reuters-21578 benchmark data and a real-world customer E-mail routing system.
Keywords
Computer Science
Machine LearningBagging predictors
16,377 Citations1996Leo Breiman
Tests on real and simulated data sets using classification and regression trees and subset selection in linear regression show that bagging can give substantial gains in accuracy.
Lecture notes in computer scienceText categorization with Support Vector Machines: Learning with many relevant features
7,925 Citations1998Thorsten Joachims
SVMs achieve substantial improvements over the currently best performing methods and behave robustly over a variety of di-erent learning tasks, eliminating the need for manual parameter tuning.
Experiments with a new boosting algorithm
7,585 Citations1996Yoav Freund, Robert E. Schapire
The Annals of StatisticsAdditive logistic regression: a statistical view of boosting (With discussion and a rejoinder by the authors)
6,842 Citations2000Jerome H. Friedman, Trevor Hastie +1 more
This work shows that this seemingly mysterious phenomenon of boosting can be understood in terms of well-known statistical principles, namely additive modeling and maximum likelihood, and develops more direct approximations and shows that they exhibit nearly identical results to boosting.
A Comparative Study on Feature Selection in Text Categorization
4,766 Citations1997Yiming Yang, Jan Pedersen
DF thresholding, the simplest method with the lowest cost in computation, can be reliably used instead of IG or CHI when the computation of these measures are too expensive, and strong correlations between the DF, IG and CHI values of a term are found.
Machine LearningBoosTexter: A Boosting-based System for Text Categorization
2,194 Citations2000Robert E. Schapire, Yoram Singer
This work describes in detail an implementation, called BoosTexter, of the new boosting algorithms for text categorization tasks, and presents results comparing the performance of Boos Texter and a number of other text-categorization algorithms on a variety of tasks.
Information RetrievalAn Evaluation of Statistical Approaches to Text Categorization
1,946 Citations1999Yiming Yang
Analysis and empirical evidence suggest that the evaluation results on some versions of Reuters were significantly affected by the inclusion of a large portion of unlabelled documents, mading those results difficult to interpret and leading to considerable confusions in the literature.
Inductive learning algorithms and representations for text categorization
1,465 Citations1998Susan Dumais, John Platt +2 more
A comparison of the effectiveness of five different automatic learning algorithms for text categorization in terms of learning speed, realtime classification speed, and classification accuracy is compared.
ACM Transactions on Information SystemsAutomated learning of decision rules for text categorization
867 Citations1994Chidanand Apté, Fred J. Damerau +1 more
It is shown that machine-generated decision rules appear comparable to human performance, while using the identical rule-based representation, and compared with other machine-learning techniques.
Feature selection and feature extraction for text categorization
578 Citations1992David Lewis
The effect of selecting varying numbers and kinds of features for use in predicting category membership was investigated on the Reuters and MUC-3 text categorization data sets and the optimal feature set size for word-based indexing was found to be surprisingly low despite the large training sets.
Predictive Data Mining: A Practical Guide
574 Citations1997Sholom M. Weiss, Nitin Indurkhya
This chapter discusses data mining, reduction and mining in the context of big data, and some of the lessons learned can be applied to other areas of science and engineering.
Future Generation Computer SystemsData mining with decision trees and decision rules
306 Citations1997Chidanand Apté, Sholom M. Weiss
This paper describes the use of decision tree and rule induction in data-mining applications and presents a synopsis of some major state-of-the-art tree andrule mining methodologies, as well as some recent advances.
Boosting and Rocchio applied to text filtering
299 Citations1998Robert E. Schapire, Yoram Singer +1 more
This paper discusses two learning algorithms for text filtering: modified Rocchio and a boosting algorithm called AdaBoost, and shows how both algorithms can be adapted to maximize any general utility matrix that associates cost for each pair of machine prediction and correct label.
arXiv (Cornell University)Mistake-Driven Learning in Text Categorization
144 Citations1997Ido Dagan, Yael Karov +1 more
This work studies three mistake-driven learning algorithms for a typical task of this nature -- text categorization and presents an algorithm, a variation of Littlestone's Winnow, which performs significantly better than any other algorithm tested on this task using a similar feature set.
