Text categorization with support vector machines
Technische Universität Dortmund Eldorado (Technische Universität Dortmund)Published 1 January 1997Open access
Thorsten Joachims
Citations507
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
Abstract
This paper explores the use of Support Vector Machines (SVMs) for learning text classifiers from examples. It analyzes the particular properties of learning with text data and identifies, why SVMs are appropriate for this task. Empirical results support the theoretical findings. SVMs achieve substantial improvements over the currently best performing methods and they behave robustly over a variety of different learning tasks. Furthermore, they are fully automatic, eliminating the need for manual parameter tuning. The paper is written in English.
Keywords
Computer Science
Machine LearningSupport-Vector Networks
33,035 Citations1995Corinna Cortes, Vladimir Vapnik
High generalization ability of support-vector networks utilizing polynomial input transformations is demonstrated and the performance of the support- vector network is compared to various classical learning algorithms that all took part in a benchmark study of Optical Character Recognition.
TechnometricsStatistical Learning Theory
26,913 Citations1999Yuhai Wu, Vladimir Vapnik
Presenting a method for determining the necessary and sufficient conditions for consistency of learning process, the author covers function estimates from small data pools, applying these estimations to real-life problems, and much more.
Foundations of statistical natural language processing
9,996 Citations1999Christopher D. Manning, Hinrich Schütze
Information Processing & ManagementTerm-weighting approaches in automatic text retrieval
9,532 Citations1988Gerard Salton, Chris Buckley
This paper summarizes the insights gained in automatic term weighting, and provides baseline single term indexing models with which other more elaborate content analysis procedures can be compared.
Program electronic library and information systemsAn algorithm for suffix stripping
8,136 Citations1980Martin Porter
An algorithm for suffix stripping is described, which has been implemented as a short, fast program in BCPL, and performs slightly better than a much more elaborate system with which it has been compared.
Lecture notes in computer scienceText categorization with Support Vector Machines: Learning with many relevant features
7,925 Citations1998Thorsten Joachims
SVMs achieve substantial improvements over the currently best performing methods and behave robustly over a variety of di-erent learning tasks, eliminating the need for manual parameter tuning.
A Comparative Study on Feature Selection in Text Categorization
4,766 Citations1997Yiming Yang, Jan Pedersen
DF thresholding, the simplest method with the lowest cost in computation, can be reliably used instead of IG or CHI when the computation of these measures are too expensive, and strong correlations between the DF, IG and CHI values of a term are found.
Machine LearningText Classification from Labeled and Unlabeled Documents using EM
2,749 Citations2000Kamal Nigam, Andrew Kachites McCallum +2 more
This paper shows that the accuracy of learned text classifiers can be improved by augmenting a small number of labeled training documents with a large pool of unlabeled documents, and presents two extensions to the algorithm that improve classification accuracy under these conditions.
Information RetrievalAn Evaluation of Statistical Approaches to Text Categorization
1,946 Citations1999Yiming Yang
Analysis and empirical evidence suggest that the evaluation results on some versions of Reuters were significantly affected by the inclusion of a large portion of unlabelled documents, mading those results difficult to interpret and leading to considerable confusions in the literature.
Inductive learning algorithms and representations for text categorization
1,465 Citations1998Susan Dumais, John Platt +2 more
A comparison of the effectiveness of five different automatic learning algorithms for text categorization in terms of learning speed, realtime classification speed, and classification accuracy is compared.
A Probabilistic Analysis of the Rocchio Algorithm with TFIDF for Text Categorization
1,264 Citations1997Thorsten Joachims
A Probabilistic analysis of the Rocchio relevance feedback algorithm, one of the most popular learning methods from information retrieval, is presented in a text categorization framework and suggests that the probabilistic algorithms are preferable to the heuristic Rocchio classifier.
Journal of Quantitative LinguisticsTowards a theory of word length distribution
93 Citations1994Gejza Wimmer, Reinhard Köhler +2 more
The compound Poisson and Ord family of distributions seems to be adequate for modeling word length distributions and the relationship of word length to other language phenomena is discussed.
The perceptron algorithm vs. Winnow
47 Citations1995Jyrki Kivinen, Manfred K. Warmuth
An adversary strategy is given that forces the Perceptron algorithm to make (N-k+1)/2 mistakes when learning k-literal disjunctions over N variables, which shows that even for simple random data, the number of mistakes made by the PerCEPTron algorithm grows almost linearly with N, even if the number k of relevant variable remains a small constant.
A freely available morphological analyzer, disambiguator and context sensitive lemmatizer for German
40 Citations1998Wolfgang Lezius, Reinhard Rapp +1 more
Morphy is an integrated tool for German morphology, part-of-speech tagging and context-sensitive lemmatization that can determine the correct root even for ambiguous word forms.
Information Processing & ManagementModelling documents with multiple poisson distributions
20 Citations1993Eugene L. Margulis
It was found that over 70% of frequently occurring terms indeed behave according to the nP distributions and the results indicate that the proportion of nP terms is even higher for the collections in which documents have similar length.
Semiotics and Computational Linguistics On Semiotic Cognitive Information Processing
9 Citations1999Burghard B. Rieger
It will be argued that fuzzy modeling allows to derive more adequate representational means whose (numerical) specificity and (procedural) definiteness may complement formats of categorial type precision and processual determinateness (which would seem cognitively inadequate).
Journal of Quantitative LinguisticsA stationary model of coherent text generation*
3 Citations1995Ju.K. Krylov
An attempt to develop the theoretical foundations of the concept of statistical arrangement of words in a coherent text on the basis of a variation model.
