Natural Language Processing and Text Mining
Published 1 January 2007
Anne Kao, Stephen R. Poteet
Citations470
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
This state-of-the-art survey is a must-have for advanced students, professionals, and researchers.
Abstract
The topic this book addresses originated from a panel discussion at the 2004 ACM SIGKDD (Special Interest Group on Knowledge Discovery and Data Mining) Conference held in Seattle, Washington, USA. We
Keywords
Computer Science
Data Mining and Knowledge DiscoveryA Tutorial on Support Vector Machines for Pattern Recognition
16,433 Citations1998Christopher J. C. Burges
There are several arguments which support the observed high accuracy of SVMs, which are reviewed and numerous examples and proofs of most of the key theorems are given.
Modern Information Retrieval
11,544 Citations1999Ricardo Baeza‐Yates, Berthier Ribeiro‐Neto
Information Processing & ManagementTerm-weighting approaches in automatic text retrieval
9,532 Citations1988Gerard Salton, Chris Buckley
This paper summarizes the insights gained in automatic term weighting, and provides baseline single term indexing models with which other more elaborate content analysis procedures can be compared.
Lecture notes in computer scienceText categorization with Support Vector Machines: Learning with many relevant features
7,925 Citations1998Thorsten Joachims
SVMs achieve substantial improvements over the currently best performing methods and behave robustly over a variety of di-erent learning tasks, eliminating the need for manual parameter tuning.
ACM Computing SurveysMachine learning in automated text categorization
7,899 Citations2002Fabrizio Sebastiani
This survey discusses the main approaches to text categorization that fall within the machine learning paradigm and discusses in detail issues pertaining to three different problems, namely, document representation, classifier construction, and classifier evaluation.
Psychological ReviewA solution to Plato's problem: The latent semantic analysis theory of acquisition, induction, and representation of knowledge.
6,039 Citations1997Thomas K. Landauer, Susan Dumais
A new general theory of acquired similarity and knowledge representation, latent semantic analysis (LSA), is presented and used to successfully simulate such learning and several other psycholinguistic phenomena.
Psychological ReviewToward a model of text comprehension and production.
5,035 Citations1978Walter Kintsch, Teun A. van Dijk
The semantic structure of texts can be described both at the local microlevel and at a more global macrolevel, and a model for text comprehension based on this notion accounts for the formation of a coherent semantic text base in terms of a cyclical process constrained by limitations of working memory.
Discourse ProcessesAn introduction to latent semantic analysis
4,817 Citations1998Thomas K. Landauer, Peter W. Foltz +1 more
The adequacy of LSA's reflection of human knowledge has been established in a variety of ways, for example, its scores overlap those of humans on standard vocabulary and subject matter tests; it mimics human word sorting and category judgments; it simulates word‐word and passage‐word lexical priming data.
A Comparative Study on Feature Selection in Text Categorization
4,766 Citations1997Yiming Yang, Jan Pedersen
DF thresholding, the simplest method with the lowest cost in computation, can be reliably used instead of IG or CHI when the computation of these measures are too expensive, and strong correlations between the DF, IG and CHI values of a term are found.
A re-examination of text categorization methods
2,660 Citations1999Yiming Yang, Xin Liu
The results show that SVM, kNN and LLSF signi cantly outperform NNet and NB when the number of positive training instances per category are small, and that all the methods perform comparably when the categories are over 300 instances.
Journal of Machine Learning ResearchRCV1: A New Benchmark Collection for Text Categorization Research
2,600 Citations2004David Lewis, Yiming Yang +2 more
This work describes the coding policy and quality control procedures used in producing the RCV1 data, the intended semantics of the hierarchical category taxonomies, and the corrections necessary to remove errorful data.
Psychological ReviewConstructing inferences during narrative text comprehension.
2,411 Citations1994Arthur C. Graesser, Murray Singer +1 more
Empirical evidence is reviewed that addresses the constructionist theory that accounts for the knowledge-based inferences that are constructed when readers comprehend narrative text and contrasts it with alternative theoretical frameworks.
The foundations of cost-sensitive learning
1,835 Citations2001Charles Elkan
It is argued that changing the balance of negative and positive training examples has little effect on the classifiers produced by standard Bayesian and decision tree learning methods, and the recommended way of applying one of these methods is to learn a classifier from the training set and then to compute optimal decisions explicitly using the probability estimates given by the classifier.
Behavior Research Methods, Instruments, & ComputersCoh-Metrix: Analysis of text on cohesion and language
1,515 Citations2004Arthur C. Graesser, Danielle S. McNamara +2 more
Standard text readability formulas scale texts on difficulty by relying on word length and sentence length, whereas Coh-Metrix is sensitive to cohesion relations, world knowledge, and language and discourse characteristics.
A maximum-entropy-inspired parser
1,496 Citations2000Eugene Charniak
A new parser for parsing down to Penn tree-bank style parse trees that achieves 90.1% average precision/recall for sentences of length 40 and less and 89.5% when trained and tested on the previously established sections of the Wall Street Journal treebank is presented.
ACM SIGKDD Explorations NewsletterMining with rarity
1,386 Citations2004Gary M. Weiss
It is demonstrated that rare classes and rare cases are very similar phenomena---both forms of rarity are shown to cause similar problems during data mining and benefit from the same remediation methods.
Learning to Classify Text Using Support Vector Machines
1,366 Citations2002Thorsten Joachims
This book gives a concise introduction to SVMs for pattern recognition, and it includes a detailed description of how to formulate text-classification tasks for machine learning.
Cognition and InstructionAre Good Texts Always Better? Interactions of Text Coherence, Background Knowledge, and Levels of Understanding in Learning From Text
1,333 Citations1996Danielle S. McNamara, Eileen Kintsch +2 more
Maximum Entropy Model for Part-Of-Speech Tagging
1,276 Citations1996Adwait Ratnaparkhi
A statistical model which trains from a corpus annotated with Part Of Speech tags and assigns them to previously unseen text with state of the art accuracy and discusses the corpus consistency problems discovered during the implementation of these features.
Journal of Memory and LanguageCausal thinking and the representation of narrative events
1,061 Citations1985Tom Trabasso, Paul van den Broek
Lexical cohesion computed by thesaural relations as an indicator of the structure of text
941 Citations1991Jane Morris, Graeme Hirst
Psychology Press eBooksLanguage Comprehension As Structure Building
890 Citations2013Morton Ann Gernsbacher
Coarse-to-fine <i>n</i>-best parsing and MaxEnt discriminative reranking
888 Citations2005Eugene Charniak, Mark Johnson
This paper describes a simple yet novel method for constructing sets of 50- best parses based on a coarse-to-fine generative parser that generates 50-best lists that are of substantially higher quality than previously obtainable.
Discourse ProcessesThe measurement of textual coherence with latent semantic analysis
750 Citations1998Peter W. Foltz, Walter Kintsch +1 more
The approach for predicting coherence through reanalyzing sets of texts from 2 studies that manipulated the coherence of texts and assessed readers’ comprehension indicates that the method is able to predict the effect of text coherence on comprehension and is more effective than simple term‐term overlap measures.
A new statistical parser based on bigram lexical dependencies
619 Citations1996Michael Collins
A new statistical parser which is based on probabilities of dependencies between head-words in the parse tree, which trains on 40,000 sentences in under 15 minutes and can be improved to over 200 sentences a minute with negligible loss in accuracy.
Multi-paragraph segmentation of expository text
558 Citations1994Marti A. Hearst
TextTiling, an algorithm for partitioning expository texts into coherent multi-paragraph discourse units which reflect the subtopic structure of the texts, is described and shown to produce segmentation that corresponds well to human judgments of the major subtopic boundaries of thirteen lengthy texts.
ACM SIGKDD Explorations NewsletterFeature selection for text categorization on imbalanced data
550 Citations2004Zhaohui Zheng, Wu Xiaoyun +1 more
This work investigates the usefulness of explicit control of that combination within a proposed feature selection framework and shows both great potential and actual merits of explicitly combining positive and negative features in a nearly optimal fashion according to the imbalanced data.
Center for the Study of Language and Information eBooksOn the coherence and structure of discourse
461 Citations1985Jerry R. Hobbs
Enhancing Supervised Learning with Unlabeled Data
459 Citations2000Sally A. Goldman, Yan Zhou
A new semi-supervised learning method called co-learning that is designed to use unlabeled data to enhance standard supervised learning algorithms to leverage off the fact that they have different representations of the hypotheses and are likely to detect different patterns in labeled data.
Building a question answering test collection
448 Citations2000Ellen M. Voorhees, Dawn M. Tice
The TREC-8 Question Answering (QA) Track was the first large-scale evaluation of domain-independent question answering systems and was used to investigate whether the evaluation methodology used for document retrieval is appropriate for a different natural language processing task.
Feature selection, perception learning, and a usability case study for text categorization
446 Citations1997Hwee Tou Ng, Wei Boon Goh +1 more
An automated learning approach to text categorization based on perception learning and a new feature selection metric, called correlation coefficient, is described and empirical results indicate that this approach outperforms the best published results on this % uters collection.
Feature Selection for Unbalanced Class Distribution and Naive Bayes
432 Citations1999Dunja Mladenić, Marko Grobelnik
This paper describes an approach to feature subset selection that takes into account problem speciics and learning algorithm characteristics, and shows that considering domain and algorithm characteristics signiicantly improves the results of classiication.
Machine LearningText Categorization with Support Vector Machines. How to Represent Texts in Input Space?
403 Citations2002Edda Leopold, Jörg Kindermann
It is shown that in the case of text classification, term-frequency transformations have a larger impact on the performance of SVM than the kernel itself.
Journal of Experimental Psychology Learning Memory and CognitionProcessing narrative time shifts.
358 Citations1996Rolf A. Zwaan
Supervised term weighting for automated text categorization
355 Citations2003Franca Debole, Fabrizio Sebastiani
It is proposed that learning from training data should also affect phase (ii), i.e. that information on the membership of training documents to categories be used to determine term weights, and is called supervised term weighting (STW).
Canadian Journal of Experimental Psychology/Revue canadienne de psychologie expérimentaleReading both high-coherence and low-coherence texts: Effects of text sequence and prior knowledge.
337 Citations2001Danielle S. McNamara
Low-knowledge readers benefited from the high-coherence text, regardless of whether it was read first, second, or twice, while high- knowledge readers benefitedfrom the low- coherence only text when it was reading first.
Journal of Educational PsychologyUsing Kintsch's computational model to improve instructional text: Effects of repairing inference calls on recall and cognitive structures.
316 Citations1991Bruce K. Britton, Sami̇ Gülgöz
Kintsch's reading comprehension model was used to identify locations where inferences were called for in a 1000-word expository text and each location was repaired to produce a principled revision.
Recognizing text genres with simple metrics using discriminant analysis
293 Citations1994Jussi Karlgren, Douglass R. Cutting
A simple method for categorizing texts into pre-determined text genre categories using the statistical standard technique of discriminant analysis is demonstrated with application to the Brown corpus.
Modern Language JournalA Linguistic Analysis of Simplified and Authentic Texts
291 Citations2007Scott A. Crossley, Max M. Louwerse +2 more
ACM SIGKDD Explorations NewsletterExtreme re-balancing for SVMs
285 Citations2004Bhavani Raskutti, Adam Kowalczyk
There is a consistent pattern of performance differences between one and two-class learning for all SVMs investigated, and these patterns persist even with aggressive dimensionality reduction through automated feature selection.
Information RetrievalHierarchical Text Categorization Using Neural Networks
248 Citations2002Miguel E. Ruiz, Padmini Srinivasan
The results show that the use of the hierarchical structure improves text categorization performance with respect to an equivalent flat model, and the optimized Rocchio algorithm achieves a performance comparable with that of the Hierarchical Mixture of Experts model.
Computers and the HumanitiesComputer-Based Authorship Attribution Without Lexical Measures
219 Citations2001Efstathios Stamatatos, Nikos Fakotakis +1 more
This paper presents a fully-automated approach to the identification of the authorship of unrestricted text that excludes any lexical measure and adapts aset of style markers to the analysis of the text performed by an already existing natural language processing tool using three stylometric levels.
Literary and Linguistic ComputingWord-Patterns and Story-Shapes: The Statistical Analysis of Narrative Style
194 Citations1987J. F. Burrows
The structure and performance of an open-domain question answering system
187 Citations2000Dan Moldovan, Sanda M. Harabagiu +5 more
The LASSO Question Answering system developed in the Natural Language Processing Laboratory at SMU relies on a combination of syntactic and semantic techniques and a novel form of indexing called paragraph indexing to find answers.
Text, speech and language technologyUnsupervised Learning of Disambiguation Rules for Part-of-Speech Tagging
186 Citations1999Eric Brill, Mihaela Pop
An unsupervised learning algorithm for automatically training a rule-based part of speech tagger without using a manually tagged corpus is described and compared to the Baum-Welch algorithm, used for unsupervised training of stochastic taggers.
Cognition and InstructionEffects of Causal Text Revisions on More- and Less-Skilled Readers' Comprehension of Easy and Difficult Texts
157 Citations2000Tracy Linderholm, Michelle Everson +4 more
Interactive Learning EnvironmentsSupporting Content-Based Feedback in On-Line Writing Evaluation with LSA
146 Citations2000Peter W. Foltz, Sara Gilliam +1 more
Tests of an automated essay grader and critic that uses Latent Semantic Analysis show that LSA can score as accurately as the humans and implications for the use of the technology in undergraduate courses and how it can provide an effective approach to incorporating more writing both in and outside of the classroom.
Interactive Learning EnvironmentsDeveloping Summarization Skills through the Use of LSA-Based Feedback
124 Citations2000Eileen Kintsch, Dave Steinhart +4 more
The collaborative process among researchers and teachers which enabled the development of a viable and supportive educational tool and its integration into classroom instruction are described.
ACM Transactions on Asian Language Information ProcessingAn adaptive <i>k</i> -nearest neighbor text categorization strategy
108 Citations2004Baoli Li, Qin Lu +1 more
Experiments show that the improved kNN strategy, in which different numbers of nearest neighbors for different categories are used instead of a fixed number across all categories, is especially applicable and promising for cases where estimating the parameter k via cross-validation is not possible and the class distribution of a training set is skewed.
Metaphor and SymbolMetaphor Comprehension: What Makes a Metaphor Difficult to Understand?
104 Citations2002Walter Kintsch, Anita R. Bowles
Journal of Educational PsychologyEffects of coherence and relevance on shallow and deep text processing.
104 Citations2002Stephen Lehman, Gregory Schraw
The hypothesis that enhancing the relevance of text segments compensates for breaks in local and global coherence was supported and one explanation for these results is that relevance enables readers to focus on salient information, which in turn can be used to repair serious coherence breaks.
Analyzing Writing Styles with Coh-Metrix.
94 Citations2006Philip M. McCarthy, Gwyneth A. Lewis +2 more
Evidence that authors within the same register can be computationally distinguished despite evidence that stylistic markers can also shift significantly over time is reported.
Reading Research QuarterlyThe Effects of Thinking Aloud during Reading on Students' Comprehension of More or Less Coherent Text
89 Citations1994Jane A. Loxterman, Isabel L. Beck +1 more
eScholarship (California Digital Library)Variation in Language and Cohesion across Written and Spoken Registers
89 Citations2004Max M. Louwerse, Philip M. McCarthy +2 more
This paper investigates the variation in cohesion across written and spoken registers and compared 236 language and cohesion features at the text- level, which showed most variation in speech and writing, whereas the linguistic feature analysis operating at the word level did not yield any difference.
Behavior Research Methods, Instruments, & ComputersUse of latent semantic analysis for predicting psychological phenomena: Two issues and proposed solutions
75 Citations2003Michael Wolfe, Susan R. Goldman
Two issues are discussed that researchers must attend to when evaluating the utility of LSA for predicting psychological phenomena, and LSA indices of similarity should be derived from theoretical analysis of the processes involved in understanding two conflicting accounts of a historical event.
Parsimonious and Profligate Approaches to the Question of Discourse Structure Relations.
65 Citations1990Eduard Hovy
This paper classifies the more than 350 relations they have proposed into a hierarchy of increasingly semantic relations, and argues that though the hierarchy is open-ended in one dimension, it is bounded in the other and therefore does not give rise to anarchy.
ACM SIGKDD Explorations NewsletterA multistrategy approach for digital text categorization from imbalanced documents
64 Citations2004M. Dolores del Castillo, J. Ignacio Serrano
A multistrategy classifier system that can be used for document categorization that relies on a modular, flexible architecture that makes no assumptions about the design of learners or the number of learners available and guarantees the independence of the thematic domain.
Lecture notes in computer scienceMining Extremely Skewed Trading Anomalies
21 Citations2004Wei Fan, Philip S. Yu +1 more
A simple systematic approach to mine very skewed distribution in very large volume of data in order to satisfy federal trading regulations as well as prevent crimes, such as insider trading and money laundry is proposed.
Handling of Imbalanced Data in Text Classification: Category-Based Term Weights
12 Citations2007Ying Liu, Han Tong Loh +2 more
Learning from imbalanced data has emerged as a new challenge to the machine learning, data mining and text mining communities and an excellent review of the state of the art is given by Gary Weiss.
DSpace@MIT (Massachusetts Institute of Technology)Building a Document Corpus for Manufacturing Knowledge Retrieval
12 Citations2004Yuhan Liu, Han Tong Loh +1 more
The origins and motivation of building MCV1, the innovative coding process which is specially designed for manufacturing companies will be presented, and all other relevant issues, like coding policy, category codes and input documents, will be explained.
Linguistic Computing with UNIX Tools
5 Citations2007Lothar M. Schmitt, Kiel Christianson +1 more
This chapter presents an outline of applications to language analysis that open up through the combined use of two simple yet powerful programming languages with particularly short descriptions: sed and awk, and demonstrates how these two UNIX1 tools can be used to implement small, useful and customized applications.
Evolving Explanatory Novel Patterns for Semantically-Based Text Mining
4 Citations2007John Atkinson
There are tools using basic pattern recognition techniques and heuristics that are capable of extracting valuable information from free text based on the elements contained in it (i.e., keywords), and this technology is usually referred to as Text Mining, and aims at discovering unseen and interesting patterns in textual databases.
