Learning extraction patterns for subjective expressions
Published 1 January 2003Open access
Ellen Riloff, Janyce Wiebe
Citations1,001
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
A bootstrapping process that learns linguistically rich extraction patterns for subjective (opinionated) expressions while maintaining high precision is presented.
Abstract
This paper presents a bootstrapping process that learns linguistically rich extraction patterns for subjective (opinionated) expressions.
Keywords
Computer Science
Thumbs up?
6,987 Citations2002Bo Pang, Lillian Lee +1 more
This work considers the problem of classifying documents not by topic, but by overall sentiment, e.g., determining whether a review is positive or negative, and concludes by examining factors that make the sentiment classification problem more challenging.
College Composition and CommunicationA Comprehensive Grammar of the English Language
5,535 Citations1987P Beauvais, Randolph Quirk +3 more
Meeting of the Association for Computational LinguisticsThumbs Up or Thumbs Down? Semantic Orientation Applied to Unsupervised Classification of Reviews
3,654 Citations2002Peter Peter, Turney
The Berkeley FrameNet Project
2,564 Citations1998Collin F. Baker, Charles J. Fillmore +1 more
This report will present the project's goals and workflow, and information about the computational tools that have been adapted or created in-house for this work.
Language<b>English verb classes and alternations:</b> A preliminary investigation. By Beth Levin. Chicago: University of Chicago Press, 1993. Pp. xviii, 345. Cloth $45.00, paper $18.95.
2,322 Citations1995Carol L. Tenny
Beth Levin shows how identifying verbs with similar syntactic behavior provides an effective means of distinguishing semantically coherent verb classes, and isolates these classes by examining verb behavior with respect to a wide range of syntactic alternations that reflect verb meaning.
arXiv (Cornell University)Thumbs up? Sentiment Classification using Machine Learning Techniques
2,208 Citations2002Bo Pang, Lillian Lee +1 more
ArXiv.orgThumbs Up or Thumbs Down? Semantic Orientation Applied to Unsupervised Classification of Reviews
1,585 Citations2002Peter D. Turney
Predicting the semantic orientation of adjectives
1,439 Citations1997Vasileios Hatzivassiloglou, Kathleen McKeown
A log-linear regression model uses constraints from conjunctions to predict whether conjoined adjectives are of same or different orientations, achieving 82% accuracy in this task when each conjunction is considered independently.
Machine LearningLearning Information Extraction Rules for Semi-Structured and Free Text
928 Citations1999Stephen Soderland
WHISK is designed to handle text styles ranging from highly structured to free text, including text that is neither rigidly formatted nor composed of grammatical sentences, and can also handle extraction from free text such as news stories.
Learning dictionaries for information extraction by multi-level bootstrapping
688 Citations1999Ellen Riloff, Rosie Jones
A multilevel bootstrapping algorithm is presented that generates both the semantic lexicon and extraction patterns simultaneously simultaneously and produces high-quality dictionaries for several semantic categories.
Automatically generating extraction patterns from untagged text
620 Citations1996Ellen Riloff
This work has developed a system called AutoSlog-TS that creates dictionaries of extraction patterns using only untagged text, and in experiments with the MUG-4 terrorism domain, created a dictionary of extraction pattern that performed comparably to a dictionary created by autoSlog, using only preclassified texts as input.
Learning Subjective Adjectives from Corpora
520 Citations2000Janyce Wiebe
This paper identifies strong clues of subjectivity using the results of a method for clustering words according to distributional similarity (Lin 1998), seeded by a small amount of detailed manual annotation.
Learning subjective nouns using extraction pattern bootstrapping
499 Citations2003Ellen Riloff, Janyce Wiebe +1 more
The goal of the research is to develop a system that can distinguish subjective sentences from objective sentences, and a Naive Bayes classifier is trained using the subjective nouns, discourse features, and subjectivity clues identified in prior research.
Development and use of a gold-standard data set for subjectivity classifications
498 Citations1999Janyce Wiebe, Rebecca Bruce +1 more
Bias-corrected tags are formulated and successfully used to guide a revision of the coding manual and develop an automatic classifier.
Automatically constructing a dictionary for information extraction tasks
453 Citations1993Ellen Riloff
Using AutoSlog, a system that automatically builds a domain-specific dictionary of concepts for extracting information from text, a dictionary for the domain of terrorist event descriptions was constructed in only 5 person-hours and the overall scores were virtually indistinguishable.
Automatic detection of text genre
358 Citations1997Brett Kessler, Geoffrey Numberg +1 more
A theory of genres as bundles of facets, which correlate with various surface cues, are proposed, and it is argued that genre detection based on surface cues is as successful as Detection based on deeper structural properties.
Recognizing text genres with simple metrics using discriminant analysis
293 Citations1994Jussi Karlgren, Douglass R. Cutting
A simple method for categorizing texts into pre-determined text genre categories using the statistical standard technique of discriminant analysis is demonstrated with application to the Brown corpus.
ArXiv.orgCRYSTAL: Inducing a Conceptual Dictionary
273 Citations1995Stephen Soderland, D A Fisher +2 more
Automatic acquisition of domain knowledge for Information Extraction
228 Citations2000Roman Yangarber, Ralph Grishman +2 more
This paper presents an alternative approach, based on an automatic discovery procedure, EXDISCO, which identifies a set of relevant documents and aSet of event patterns from un-annotaled text, starting from a small set of "seed patterns."
arXiv (Cornell University)Automatic Detection of Text Genre
206 Citations1997Brett Kessler, Geoffrey Nunberg +1 more
Linguistic VariationUnspeakable sentences
195 Citations2017Liliane Haegeman
The paper develops the cartographic analysis of register based subject omission proposed in Haegeman (2013) and based on Rizzi’s (2006b) ‘Privilege of the Root’ approach, hypothesis that there is a specialized projection for the encoding of subjecthood (SubjP).
Smokey: automatic recognition of hostile messages
170 Citations1997Ellen Spertus
Some approaches to flame recognition are described, mcluding a prototype system, Smokey, which builds a 47-element feature vector based on the syntax and semantics of each sentence, combining the vectors for the sentences within each message.
Lecture notes in computer scienceLearning information extraction patterns from examples
146 Citations1996Scott B. Huffman
A system that can learn dictionaries of extraction patterns directly from user-provided examples of texts and events to be extracted from them, and learns patterns that recognize relationships between key constituents based on local syntax.
Identifying Collocations for Recognizing Opinions
135 Citations2001Jens Wiebe
Promising results are shown for a straightforward method of identifying collocational clues of subjectivity, as well as evidence of the usefulness of these clues for recognizing opinionated documents.
Language<b>Speech act classification</b> : A study in the lexical analysis of English speech activity verbs. By Thomas Ballmer and Waltraud Brennenstuhl. (Springer series in language and communication, 8.) Berlin: Springer, 1981. Pp. x, 274.
112 Citations1983Jef Verschueren
The author's Motivation for a Speech Act Classification is explained and a survey of the Resulting Speech Act classification is conducted.
Relational learning techniques for natural language information extraction
88 Citations1998Mary Elaine Califf, Raymond J. Mooney
Experimental results show that the number of examples required to achieve a given level of performance can be significantly reduced by this method, demonstrating the superiority of relational learning for some information extraction tasks.
Annotating Opinions in the World Press
83 Citations2003Theresa Wilson, Janyce Wiebe
This paper presents a detailed scheme for annotating expressions of opinions, beliefs, emotions, sentiment and speculation in the news and other discourse, and explores inter-annotator agreement for individual private state expressions.
Toward general-purpose learning for information extraction
77 Citations1998Dayne Freitag
SRV is described, a learning architecture for information extraction which is designed for maximum generality and flexibility and can exploit domain-specific information, including linguistic syntax and lexical information, in the form of features provided to the system explicitly as input for training.
Acquisition of semantic patterns for information extraction from corpora
69 Citations2002J.-T. Kim, Dan Moldovan
A knowledge acquisition tool to extract semantic patterns for a memory-based information retrieval system is presented to facilitate the construction of a large knowledge base of semantic patterns.
Recognizing subjective sentences: a computational investigation of narrative text
53 Citations1990Janyce Wiebe
It is shown that references are understood differently in conversation, subjective sentences, and non-subjective sentences; namely, they are understood with respect to different sets of beliefs.
Proceedings of the 17th international conference on Computational linguistics -Toward general-purpose learning for information extraction
31 Citations1998Dayne Freitag
