Supervised term weighting for automated text categorization
Published 9 March 2003Open access
Franca Debole, Fabrizio Sebastiani
Citations355
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
It is proposed that learning from training data should also affect phase (ii), i.e. that information on the membership of training documents to categories be used to determine term weights, and is called supervised term weighting (STW).
Abstract
Researchers from ISTI-CNR, Pisa, aim at producing better text classification methods through the use of supervised learning techniques in the generation of the internal representations of the texts
Keywords
Computer Science
Elements of Information Theory
37,533 Citations2001Thomas M. Cover, Joy A. Thomas
Machine LearningInduction of Decision Trees
14,815 Citations1986J. R. Quinlan
This paper summarizes an approach to synthesizing decision trees that has been used in a variety of systems, and it describes one such system, ID3, in detail, which is described in detail.
Foundations of statistical natural language processing
9,996 Citations1999Christopher D. Manning, Hinrich Schütze
Information Processing & ManagementTerm-weighting approaches in automatic text retrieval
9,532 Citations1988Gerard Salton, Chris Buckley
This paper summarizes the insights gained in automatic term weighting, and provides baseline single term indexing models with which other more elaborate content analysis procedures can be compared.
ACM Computing SurveysMachine learning in automated text categorization
7,899 Citations2002Fabrizio Sebastiani
This survey discusses the main approaches to text categorization that fall within the machine learning paradigm and discusses in detail issues pertaining to three different problems, namely, document representation, classifier construction, and classifier evaluation.
International Conference on Neural Information ProcessingAdvances in kernel methods: support vector learning
5,815 Citations1999Bernhard Schölkopf, Christopher J. C. Burges +1 more
Support vector machines for dynamic reconstruction of a chaotic system, Klaus-Robert Muller et al pairwise classification and support vector machines, Ulrich Kressel.
A Comparative Study on Feature Selection in Text Categorization
4,766 Citations1997Yiming Yang, Jan Pedersen
DF thresholding, the simplest method with the lowest cost in computation, can be reliably used instead of IG or CHI when the computation of these measures are too expensive, and strong correlations between the DF, IG and CHI values of a term are found.
Technical reportsMaking Large-Scale SVM Learning Practical
4,317 Citations2006Thorsten Joachims
This chapter presents algorithmic and computational results developed for SVM light V 2.0, which make large-scale SVM training more practical and give guidelines for the application of SVMs to large domains.
A re-examination of text categorization methods
2,660 Citations1999Yiming Yang, Xin Liu
The results show that SVM, kNN and LLSF signi cantly outperform NNet and NB when the number of positive training instances per category are small, and that all the methods perform comparably when the categories are over 300 instances.
Lecture notes in computer scienceNaive (Bayes) at forty: The independence assumption in information retrieval
2,092 Citations1998David Lewis
The naive Bayes classifier, currently experiencing a renaissance in machine learning, has long been a core technique in information retrieval, and some of the variations used for text retrieval and classification are reviewed.
Inductive learning algorithms and representations for text categorization
1,465 Citations1998Susan Dumais, John Platt +2 more
A comparison of the effectiveness of five different automatic learning algorithms for text categorization in terms of learning speed, realtime classification speed, and classification accuracy is compared.
Morgan Kaufmann Publishers Inc. eBooksReadings in information retrieval
1,064 Citations1997Karen Spärck Jones, Peter Willett
Induction of Decision Trees
999 Citations2003Quinlan
Feature Selection for Unbalanced Class Distribution and Naive Bayes
432 Citations1999Dunja Mladenić, Marko Grobelnik
This paper describes an approach to feature subset selection that takes into account problem speciics and learning algorithm characteristics, and shows that considering domain and algorithm characteristics signiicantly improves the results of classiication.
ACM SIGIR ForumExploring the similarity space
386 Citations1998Justin Zobel, Alistair Moffat
It is demonstrated that it is surprisingly difficult to identify which techniques work best, and comment on the experimental methodology required to support any claims as to the superiority of one method over another.
Scholarworks (University of Massachusetts Amherst)Representation and Learning in Information Retrieval
361 Citations1991David Lewis
A new theoretical model for text classification systems, including systems for document retrieval, automated indexing, electronic mail filtering, and similar tasks, is introduced, suggesting that the poor statistical characteristics of a syntactic indexing phrase representation negate its desirable semantic characteristics.
Evaluating and optimizing autonomous text classification systems
312 Citations1995David Lewis
This work shows how to define what constitutes good effectiveness for binary text classification systems, tune the systems to achieve the highest possible effectiveness, and estimate how the effectiveness changes as new data is processed.
Lecture notes in computer scienceExperiments on the Use of Feature Selection and Negative Evidence in Automated Text Categorization
209 Citations2000Luigi Galavotti, Fabrizio Sebastiani +1 more
This work proposes a novel variant, based on the exploitation of negative evidence, of the well-known k-NN method, and reports the results of systematic experimentation of these two methods performed on the standard REUTERS-21578 benchmark.
Feature selection in SVM text categorization
108 Citations1999Hirotoshi Taira, Masahiko Haruno
Results suggest a simple strategy for the SVM text categorization: use a full number of words found through a rough filtering technique like part-of-speech tagging, which indicates that SVMs cannot find irrelevant parts of speech.
