The text mining handbook: advanced approaches in analyzing unstructured data
Choice Reviews OnlinePublished 1 June 2007
Citations1,626
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
Abstract
1. Introduction to text mining 2. Core text mining operations 3. Text mining preprocessing techniques 4. Categorization 5. Clustering 6. Information extraction 7. Probabilistic models for Information extraction 8. Preprocessing applications using probabilistic and hybrid approaches 9. Presentation-layer considerations for browsing and query refinement 10. Visualization approaches 11. Link analysis 12. Text mining applications Appendix Bibliography.
Keywords
Computer Science
The Nature of Statistical Learning Theory
39,279 Citations1995Vladimir Vapnik
Elements of Information Theory
37,533 Citations2001Thomas M. Cover, Joy A. Thomas
Journal of Machine Learning ResearchLatent dirichlet allocation
27,049 Citations2003David M. Blei, Andrew Y. Ng +1 more
Social NetworksCentrality in social networks conceptual clarification
16,883 Citations1978Linton C. Freeman
Three distinct intuitive conceptions of centrality are uncovered and existing measures are refined to embody these conceptions, one absolute and one relative measure of the centrality of positions in a network and one reflecting the degree of centralization of the entire network.
Journal of the American Society for Information ScienceIndexing by latent semantic analysis
12,677 Citations1990Scott Deerwester, Susan Dumais +3 more
Modern Information Retrieval
11,544 Citations1999Ricardo Baeza‐Yates, Berthier Ribeiro‐Neto
Lecture notes in computer scienceText categorization with Support Vector Machines: Learning with many relevant features
7,925 Citations1998Thorsten Joachims
SVMs achieve substantial improvements over the currently best performing methods and behave robustly over a variety of di-erent learning tasks, eliminating the need for manual parameter tuning.
ACM Computing SurveysMachine learning in automated text categorization
7,899 Citations2002Fabrizio Sebastiani
This survey discusses the main approaches to text categorization that fall within the machine learning paradigm and discusses in detail issues pertaining to three different problems, namely, document representation, classifier construction, and classifier evaluation.
Building a Large Annotated Corpus of English: The Penn Treebank
7,528 Citations1993Mitchell P. Marcus
Software Practice and ExperienceGraph drawing by force‐directed placement
6,359 Citations1991Thomas M. J. Fruchterman, Edward M. Reingold
A modification of the spring‐embedder model of Eades for drawing undirected graphs with straight edges is presented, developed in analogy to forces in natural systems, for a simple, elegant, conceptually‐intuitive, and efficient algorithm.
Contemporary Sociology A Journal of ReviewsSocial Network Analysis: A Handbook.
6,046 Citations1993Franz Urban Pappi, James G. Scott
American Journal of SociologyPower and Centrality: A Family of Measures
5,205 Citations1987Phillip Bonacich
General Pharmacology The Vascular SystemThe visual display of quantitative information
5,120 Citations1985
Geocarto InternationalIntroductory digital image processing: A remote sensing perspective
5,040 Citations1987John R. Jensen, Kalmesh Lulla
This text focuses exclusively on the art and science of digital image processing of satellite and aircraft-derived remotely-sensed data for resource management and explains how to extract biophysical information from remote sensor data for almost all multidisciplinary land-based environmental projects.
Discourse ProcessesAn introduction to latent semantic analysis
4,817 Citations1998Thomas K. Landauer, Peter W. Foltz +1 more
The adequacy of LSA's reflection of human knowledge has been established in a variety of ways, for example, its scores overlap those of humans on standard vocabulary and subject matter tests; it mimics human word sorting and category judgments; it simulates word‐word and passage‐word lexical priming data.
Medical Entomology and ZoologyHead-driven phrase structure grammar
3,973 Citations1994Ivan A. Sag, Carl Pollard
This book presents the most complete exposition of the theory of head-driven phrase structure grammar, introduced in the authors' "Information-Based Syntax and Semantics," and demonstrates the applicability of the HPSG approach to a wide range of empirical problems.
Readings in Information Visualization: Using Vision to Think
3,967 Citations1999Stuart K. Card, Jock D. Mackinlay +1 more
Journal of Mathematical SociologyFactoring and weighting approaches to status scores and clique identification
3,129 Citations1972Phillip Bonacich
Machine LearningOn the Optimality of the Simple Bayesian Classifier under Zero-One Loss
3,072 Citations1997Pedro Domingos, Michael J. Pazzani
The Bayesian classifier is shown to be optimal for learning conjunctions and disjunctions, even though they violate the independence assumption, and will often outperform more powerful classifiers for common training set sizes and numbers of attributes, even if its bias is a priori much less appropriate to the domain.
A re-examination of text categorization methods
2,660 Citations1999Yiming Yang, Xin Liu
The results show that SVM, kNN and LLSF signi cantly outperform NNet and NB when the number of positive training instances per category are small, and that all the methods perform comparably when the categories are over 300 instances.
Machine LearningBoosTexter: A Boosting-based System for Text Categorization
2,194 Citations2000Robert E. Schapire, Yoram Singer
This work describes in detail an implementation, called BoosTexter, of the new boosting algorithms for text categorization tasks, and presents results comparing the performance of Boos Texter and a number of other text-categorization algorithms on a variety of tasks.
Series on software engineering and knowledge engineeringData Structures and Algorithms
2,131 Citations2003
The basis of this book is the material contained in the first six chapters of the earlier work, The Design and Analysis of Computer Algorithms, and has added material on algorithms for external storage and memory management.
Lecture notes in computer scienceNaive (Bayes) at forty: The independence assumption in information retrieval
2,092 Citations1998David Lewis
The naive Bayes classifier, currently experiencing a renaissance in machine learning, has long been a core technique in information retrieval, and some of the variations used for text retrieval and classification are reviewed.
Cambridge University Press eBooksExploratory Social Network Analysis with Pajek
2,083 Citations2018Wouter de Nooy, Andrej Mrvar +1 more
Elsevier eBooksNewsWeeder: Learning to Filter Netnews
2,022 Citations1995Ken Lang
The results show that a learning algorithm based on the Minimum Description Length (MDL) principle was able to raise the percentage of interesting articles to be shown to users from 14% to 52% on average.
Information RetrievalAn Evaluation of Statistical Approaches to Text Categorization
1,946 Citations1999Yiming Yang
Analysis and empirical evidence suggest that the evaluation results on some versions of Reuters were significantly affected by the inclusion of a large portion of unlabelled documents, mading those results difficult to interpret and leading to considerable confusions in the literature.
Social NetworksNetwork structure and minimum degree
1,873 Citations1983Stephen B. Seidman
An approach to network cohesion is proposed that is based on minimum degree and which produces a sequence of subgraphs of gradually increasing cohesion that associates with any network measures of local density which promise to be useful both in characterizing network structures and in comparing networks.
Journal of Mathematical SociologyStructural equivalence of individuals in social networks
1,689 Citations1971François Lorrain, Harrison C. White
Transformation-based error-driven learning and natural language processing: a case study in part-of-speech tagging
1,531 Citations1995Eric Brill
This paper describes a simple rule-based approach to automated learning of linguistic knowledge that has been shown for a number of tasks to capture information in a clearer and more direct fashion without a compromise in performance.
Choice Reviews OnlineInformation seeking in electronic environments
1,494 Citations1996Gary Marchionini
Wrapper induction for information extraction
1,044 Citations1997Nicholas Kushmerick, Daniel S. Weld
This work introduces wrapper induction, a method for automatically constructing wrappers, and identifies hlrt, a wrapper class that is e(cid:14)ciently learnable, yet expressive enough to handle 48% of a recently surveyed sample of Internet resources.
Tackling the poor assumptions of naive bayes text classifiers
952 Citations2003Jason D. M. Rennie, Lawrence Shih +2 more
This paper proposes simple, heuristic solutions to some of the problems with Naive Bayes classifiers, addressing both systemic issues as well as problems that arise because text is not actually generated according to a multinomial model.
ACM Transactions on Information SystemsAutomated learning of decision rules for text categorization
867 Citations1994Chidanand Apté, Fred J. Damerau +1 more
It is shown that machine-generated decision rules appear comparable to human performance, while using the identical rule-based representation, and compared with other machine-learning techniques.
Hierarchically Classifying Documents Using Very Few Words
840 Citations1997Daphne Koller, Mehran Sahami
This work proposes an approach that utilizes the hierarchical topic structure to decompose the classification task into a set of simpler problems, one at each node in the classification tree, which can be solved accurately by focusing only on a very small set of features, those relevant to the task at hand.
Information RetrievalLearning Algorithms for Keyphrase Extraction
837 Citations2000Peter D. Turney
The experimental results support the claim that a custom-designed algorithm (GenEx), incorporating specialized procedural domain knowledge, can generate better keyphrases than a general-purpose algorithm (C4.5).
Hierarchical classification of Web content
805 Citations2000Susan Dumais, Hao Chen
This paper explores the use of hierarchical structure for classifying a large, heterogeneous collection of web content using support vector machine (SVM) classifiers, which have been shown to be efficient and effective for classification, but not previously explored in the context of hierarchical classification.
Machine LearningAn Algorithm that Learns What's in a Name
784 Citations1999Daniel M. Bikel, Richard Schwartz +1 more
IdentiFinderTM, a hidden Markov model that learns to recognize and classify names, dates, times, and numerical quantities, is evaluated and is competitive with approaches based on handcrafted rules on mixed case text and superior on text where case information is not available.
Computational LinguisticsAn algorithm for pronominal anaphora resolution
758 Citations1994Shalom Lappin, Herbert J. Leass
The algorithm applies to the syntactic representations generated by McCord's Slot Grammar parser and relies on salience measures derived from syntactic structure and a simple dynamic model of attentional state and to models of anaphora resolution that invoke a variety of informational factors in ranking antecedent candidates.
Communications of the ACMInformation extraction
726 Citations1996Jim Cowie, Wendy G. Lehnert
A relatively new development—information extraction (IE)—is the subject of this article and can transform the raw material, refining and reducing it to a germ of the original text.
Lecture notes in computer sciencePajek— Analysis and Visualization of Large Networks
717 Citations2002Vladimir Batagelj, Andrej Mrvar
Computer NetworksGrouper: a dynamic clustering interface to Web search results
716 Citations1999Oren Zamir, Oren Etzioni
This paper introduces Grouper, an interface to the results of the HuskySearch meta-search engine, which dynamically groups the search results into clusters labeled by phrases extracted from the snippets, and reports on the first empirical comparison of user Web search behavior on a standard ranked-list presentation versus a clustered presentation.
Machine LearningStatistical Models for Text Segmentation
687 Citations1999Doug Beeferman, Adam Berger +1 more
Assessment of the approach on quantitative and qualitative grounds demonstrates its effectiveness in two very different domains, Wall Street Journal news articles and television broadcast news story transcripts, using a new probabilistically motivated error metric.
Learning to extract symbolic knowledge from the World Wide Web
675 Citations1998Mark Craven, Dan DiPasquo +5 more
The goal of the research described here is to automatically create a computer understandable world wide knowledge base whose content mirrors that of the World Wide Web, and several machine learning algorithms for this task are described.
IEEE Transactions on Software EngineeringA technique for drawing directed graphs
668 Citations1993Emden R. Gansner, E. Koutsofios +2 more
A four-pass algorithm for drawing directed graphs is presented, which creates good drawings and is fast.
Literary and Linguistic ComputingAutomatically Categorizing Written Texts by Author Gender
633 Citations2002Moshe Koppel
It is shown that automated text categorization techniques can exploit combinations of simple lexical and syntactic features to infer the gender of the author of an unseen formal written document with approximately 80 per cent accuracy.
University of Minnesota Digital Conservancy (University of Minnesota)Criterion Functions for Document Clustering: Experiments and Analysis
592 Citations2001Ying Zhao, George Karypis
Question classification using support vector machines
582 Citations2003Dell Zhang, Wee Sun Lee
This paper proposes to use a special kernel function called the tree kernel to enable the SVM to take advantage of the syntactic structures of questions, and describes how the tree Kernel can be computed efficiently by dynamic programming.
Artificial IntelligenceWrapper induction: Efficiency and expressiveness
558 Citations2000Nicholas Kushmerick
This article describes six wrapper classes, and uses a combination of empirical and analytical techniques to evaluate the computational tradeoffs among them, finding that most of their wrapper classes are reasonably useful, yet can rapidly learned.
Journal of the ACMAutomatic Indexing: An Experimental Inquiry
552 Citations1961M. E. Maron
The design, execution and evaluation of a modest experimental study aimed at testing empirically one statistical technique for automatic indexed documents according to their subject content are described.
An evaluation of phrasal and clustered representations on a text categorization task
547 Citations1992David Lewis
It is shown that optimal effectiveness occurs when using only a small proportion of the indexing terms available, and that effectiveness peaks at a higher feature set size and lower effectiveness level for a syntactic phrase indexing than for word-based indexing.
Computers and the HumanitiesA method for disambiguating word senses in a large corpus
534 Citations1992William A. Gale, Kenneth Church +1 more
The proposed method was designed to disambiguate senses that are usually associated with different topics using a Bayesian argument that has been applied successfully in related tasks such as author identification and information retrieval.
Handbook of Data Mining and Knowledge Discovery
518 Citations2002Willi Klösgen, Jan M. Żytkow
ComputerMining the Web's link structure
502 Citations1999Soumen Chakrabarti, Byron Dom +6 more
Clever is a search engine that analyzes hyperlinks to uncover two types of pages: authorities, which provide the best source of information on a given topic; and hubs, which provides collections of links to authorities.
NeurocomputingWEBSOM – Self-organizing maps of document collections
493 Citations1998Samuel Kaski, Timo Honkela +2 more
Special consideration is given to the computation of very large document maps which is possible with general-purpose computers if the dimensionality of the word category histograms is first reduced with a random mapping method and if computationally efficient algorithms are used in computing the SOMs.
ACM Transactions on GraphicsDrawing graphs nicely using simulated annealing
489 Citations1996Ron Davidson, David Harel
The paradigm of simulated annealing is applied to the problem of drawing graphs “nicely,” and the algorithm deals with general undirected graphs with straight-line edges, and employs several simple criteria for the aesthetic quality of the result.
ACM SIGMOD RecordMining e-mail content for author identification forensics
485 Citations2001Olivier De Vel, Alison Anderson +2 more
An investigation into e-mail content mining for author identification, or authorship attribution, for the purpose of forensic investigation found promising results for both aggregated and multi-topic author categorisation.
Artificial IntelligenceLearning to construct knowledge bases from the World Wide Web
472 Citations2000Mark Craven, Dan DiPasquo +5 more
The goal of the research described here is to automatically create a computer understandable knowledge base whose content mirrors that of the World Wide Web, and several machine learning algorithms for this task are described, and promising initial results with a prototype system that has created a knowledge base describing university people, courses, and research projects.
A comparison of classifiers and document representations for the routing problem
451 Citations1995Hinrich Schütze, David A. Hull +1 more
This paper considers three classification techniques which have decision rules that are derived via explicit error minimization linear discriminant analysis, logistic regression, and neuraf networks, and finds that features based on latent semantic indexing are more effective for techniques such aslinear discriminant anaf-ysis and logistic regressors, which have no way to protect against overfitting.
Feature selection, perception learning, and a usability case study for text categorization
446 Citations1997Hwee Tou Ng, Wei Boon Goh +1 more
An automated learning approach to text categorization based on perception learning and a new feature selection metric, called correlation coefficient, is described and empirical results indicate that this approach outperforms the best published results on this % uters collection.
Data Mining and Knowledge DiscoveryThe Role of Occam's Razor in Knowledge Discovery
440 Citations1999Pedro Domingos
It is argued that Occam's razor's continued use in KDD risks causing significant opportunities to be missed, and should therefore be restricted to the comparatively few applications where it is appropriate.
Feature Selection for Unbalanced Class Distribution and Naive Bayes
432 Citations1999Dunja Mladenić, Marko Grobelnik
This paper describes an approach to feature subset selection that takes into account problem speciics and learning algorithm characteristics, and shows that considering domain and algorithm characteristics signiicantly improves the results of classiication.
Detecting Concept Drift with Support Vector Machines
413 Citations2000Ralf Klinkenberg, Thorsten Joachims
A new method to recognize and handle concept changes with support vector machines is proposed and it is shown that it can e(cid:11)ectively select an appropriate window size in a robust way.
ACM Transactions on Information SystemsAn example-based mapping method for text categorization and retrieval
405 Citations1994Yiming Yang, Christopher G. Chute
It is evident that the LLSF approach uses the relevance information effectively within human decisions of categorization and retrieval, and achieves a semantic mapping of free texts to their representations in an indexing language.
Computational LinguisticsAutomatic Text Categorization in Terms of Genre and Author
401 Citations2000Efstathios Stamatatos, Nikos Fakotakis +1 more
This paper proposes a set of style markers including analysis-level measures that represent the way in which the input text has been analyzed and capture useful stylistic information without additional cost to take full advantage of existing natural language processing (NLP) tools.
ACM SIGIR ForumA sequential algorithm for training text classifiers
382 Citations1995David Lewis
A bug in my experimental software caused the relevance sampling results reported in the SIGIR '94 paper to be incorrect, and this note presents the corrected results, along with additional data supporting the original claim that uncertainty sampling has an advantage over relevance sampling in most training situations.
Information RetrievalText Categorization Based on Regularized Linear Classification Methods
378 Citations2001Tong Zhang, Frank J. Oles
A number of known linear classification methods as well as some variants in the framework of regularized linear systems are compared to discuss the statistical and numerical properties of these algorithms, with a focus on text categorization.
Data Mining and Knowledge DiscoveryBeyond Market Baskets: Generalizing Association Rules to Dependence Rules
374 Citations1998Craig Silverstein, Sergey Brin +1 more
This work develops the notion of dependence rules that identify statistical dependence in both the presence and absence of items in itemsets in the lattice and develops pruning strategies based on the closure property that lead to an efficient algorithm for discovering dependence rules.
Journal of the American Society for Information ScienceStemming algorithms: A case study for detailed evaluation
371 Citations1996David A. Hull
A case study of stemming algorithms is described which describes a number of novel approaches to evaluation and demonstrates their value.
Scholarworks (University of Massachusetts Amherst)Representation and Learning in Information Retrieval
361 Citations1991David Lewis
A new theoretical model for text classification systems, including systems for document retrieval, automated indexing, electronic mail filtering, and similar tasks, is introduced, suggesting that the poor statistical characteristics of a syntactic indexing phrase representation negate its desirable semantic characteristics.
Bringing order to the Web
360 Citations2000Hao Chen, Susan Dumais
A user interface that organizes Web search results into hierarchical categories that allows users to focus on items in categories of interest rather than having to browse through all the results sequentially.
ACM Transactions on Information SystemsContext-sensitive learning methods for text categorization
359 Citations1999William W. Cohen, Yoram Singer
RIPPER and sleeping-experts perform extremely well across a wide variety of categorization problems, generally outperforming previously applied learning methods and are viewed as a confirmation of the usefulness of classifiers that represent contextual information.
Supervised term weighting for automated text categorization
355 Citations2003Franca Debole, Fabrizio Sebastiani
It is proposed that learning from training data should also affect phase (ii), i.e. that information on the membership of training documents to categories be used to determine term weights, and is called supervised term weighting (STW).
Machine LearningMachine Learning for Information Extraction in Informal Domains
355 Citations2000Dayne Freitag
A multistrategy approach which combines these learners and yields performance competitive with or better than the best of them is described, which is modular and flexible, and could find application in other machine learning problems.
A study of thresholding strategies for text categorization
354 Citations2001Yiming Yang
Experimental results show that the choice of thresholding strategy can significantly influence the performance of kNN, and that the ``optimal'' strategy may vary by application.
Lecture notes in computer scienceText Categorization Using Weight Adjusted k-Nearest Neighbor Classification
353 Citations2001Eui-Hong Han, George Karypis +1 more
A Weight Adjusted k-Nearest Neighbor (WAKNN) classification that learns feature weights based on a greedy hill climbing technique and two performance optimizations of WAKNN that improve the computational performance by a few orders of magnitude, but do not compromise on the classification quality.
Applied IntelligenceAuthorship Attribution with Support Vector Machines
339 Citations2003Joachim Diederich, Jörg Kindermann +2 more
The support vector machine (SVM) is applied to the use of text-mining methods for the identification of the author of a text, as it is able to cope with half a million of inputs it requires no feature selection and can process the frequency vector of all words of atext.
Learning to classify text from labeled and unlabeled documents
330 Citations1998Kamal Nigam, Andrew McCallum +2 more
It is shown that the accuracy of text classifiers trained with a small number of labeled documents can be improved by augmenting this small training set with a large pool of unlabeled documents, and an algorithm is introduced based on the combination of Expectation-Maximization with a naive Bayes classifier.
Technische Universität Dortmund Eldorado (Technische Universität Dortmund)Estimating the generalization performance of a SVM efficiently
328 Citations1999Thorsten Joachims
Without any computation-intensive resampling, the new estimators developed here are computationally much more e cient than cross-validation or bootstrapping and address the special performancemeasures needed for evaluating text classi ers.
Evaluating and optimizing autonomous text classification systems
312 Citations1995David Lewis
This work shows how to define what constitutes good effectiveness for binary text classification systems, tune the systems to achieve the highest possible effectiveness, and estimate how the effectiveness changes as new data is processed.
Information Processing & ManagementThe use of bigrams to enhance text categorization
290 Citations2002Chade-Meng Tan, Yuan-Fang Wang +1 more
An efficient text categorization algorithm that generates bigrams selectively by looking for ones that have an especially good chance of being useful by using the information gain metric, combined with various frequency thresholds is presented.
Applied Artificial IntelligenceRough set-aided keyword reduction for text categorization
266 Citations2001Alexios Chouchoulas, Qiang Shen
This article investigates the applicability of RS theory to the IF/IR application domain and compares this applicability with respect to various existing TC techniques, and investigates the ability of the approach to generalize, given a minimum of training data.
Coping with ambiguity and unknown words through probabilistic models
261 Citations1993Ralph Weischedel, Richard Schwartz +3 more
Using web structure for classifying and describing web pages
256 Citations2002Eric J. Glover, Kostas Tsioutsiouliklis +3 more
By ranking words and phrases in the citing documents according to expected entropy loss, this work is able to accurately name clusters of web pages, even with very few positive examples.
arXiv (Cornell University)FASTUS: A Cascaded Finite-State Transducer for Extracting Information from Natural-Language Text
255 Citations1997Jerry R. Hobbs, Douglas E. Appelt +5 more
This decomposition of language processing enables the system to do exactly the right amount of domain-independent syntax, so that domain-dependent semantic and pragmatic processing can be applied to the right larger-scale structures.
ACM Transactions on Computer-Human InteractionNavigation patterns and usability of zoomable user interfaces with and without an overview
237 Citations2002Kasper Hornbæk, Benjamin B. Bederson +1 more
No difference between interfaces in subjects' ability to solve tasks correctly is found, and subjects who switched between the overview and the detail windows used more time, suggesting that integration of overview and detail windows adds complexity and requires additional mental and motor effort.
ACM SIGIR ForumAutomated categorization in the international patent classification
230 Citations2003C. J. Fall, Attila Törcsvári +2 more
This work investigates how best to resolve the training problems related to the attribution of multiple classification codes to each patent document and reports the results of applying a variety of machine learning algorithms to the automated categorization of English-language patent documents.
Journal of Intelligent Information SystemsLatent Semantic Kernels
227 Citations2002Nello Cristianini, John Shawe‐Taylor +1 more
This paper describes how the LSI approach can be implemented in a kernel-defined feature space and provides experimental results demonstrating that the approach can significantly improve performance, and that it does not impair it.
A simple KNN algorithm for text categorization
225 Citations2002Pascal Soucy, Guy W. Mineau
The KNN algorithm that is proposed becomes efficient for classifying text documents in that context (in terms of its predictability and interpretability), as is demonstrated.
Proceedings of International Conference on Neural Networks (ICNN'97)Exploration of very large databases by self-organizing maps
219 Citations2002Teuvo Kohonen
A data organization system and genuine content-addressable memory called the WEBSOM, a two-layer self-organizing map (SOM) architecture where documents become mapped as points on the upper map, in a geometric order that describes the similarity of their contents.
Knowledge DIscovery in Databases:An Overview
216 Citations1991William Frawley, Gregory Piatetsky-Shapiro +1 more
Journal of the American Society for Information ScienceMap displays for information retrieval
216 Citations1997Xia Lin
A map display generated by a neural network's self-organizing algorithm detects complex relationships among given documents, and reveals the relationships through a spatial arrangement of terms abstracted from the documents.
CERN Document Server (European Organization for Nuclear Research)The Handbook of Data Mining
216 Citations2003Nong Ye
Using Reinforcement Learning to Spider the Web Efficiently
213 Citations1999Jason D. M. Rennie, Andrew McCallum
This paper presents an algorithm for learning a value function that maps hyperlinks to future discounted reward using a naive Bayes text classifier and shows a threefold improvement in spidering efficiency over traditional breadth-first search, and up to a two-fold improvement over reinforcement learning with immediate reward.
Lecture notes in computer scienceExperiments on the Use of Feature Selection and Negative Evidence in Automated Text Categorization
209 Citations2000Luigi Galavotti, Fabrizio Sebastiani +1 more
This work proposes a novel variant, based on the exploitation of negative evidence, of the well-known k-NN method, and reports the results of systematic experimentation of these two methods performed on the standard REUTERS-21578 benchmark.
arXiv (Cornell University)Automatic Detection of Text Genre
206 Citations1997Brett Kessler, Geoffrey Nunberg +1 more
…
