Identifying Collocations for Recognizing Opinions
Published 1 January 2001
Jens Wiebe
Citations135
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
Promising results are shown for a straightforward method of identifying collocational clues of subjectivity, as well as evidence of the usefulness of these clues for recognizing opinionated documents.
Abstract
Subjectivity in natural language refers to aspects of language used to express opinions and evaluations (Banfield, 1982
Keywords
Computer Science
Building a Large Annotated Corpus of English: The Penn Treebank
7,528 Citations1993Mitchell P. Marcus
Longman dictionary of contemporary English
2,421 Citations1978Paul Procter
The Longman dictionary of contemporary English is a collection ofverbs, idioms andverbs used in English since the mid-19th century that reflect the changing nature of the language.
Journal of the Royal Statistical Society Series C (Applied Statistics)Maximum Likelihood Estimation of Observer Error-Rates Using the EM Algorithm
1,666 Citations1979A. P. Dawid, A. M. Skene
The EM algorithm is shown to provide a slow but sure way of obtaining maximum likelihood estimates of the parameters of interest in compiling a patient record.
BiometrikaExploratory latent structure analysis using both identifiable and unidentifiable models
1,585 Citations1974Leo A. Goodman
Automatic retrieval and clustering of similar words
1,555 Citations1998Dekang Lin
A word similarity measure based on the distributional pattern of words allows the automatically constructed thesaurus to be significantly closer to WordNet than Roget Thesaurus is.
A simple rule-based part of speech tagger
1,418 Citations1992Eric Brill
This work presents a simple rule-based part of speech tagger which automatically acquires its rules and tags with accuracy comparable to stochastic taggers, demonstrating that the stochastics method is not the only viable method for part ofspeech tagging.
News as Discourse
1,168 Citations1988Teun Adrianus van Dijk
Retrieving collocations from text: Xtract
799 Citations1993Frank Smadja
A set of techniques based on statistical methods for retrieving and identifying collocations from large textual corpora, based on some original filtering methods that allow the production of richer and higher-precision output are described.
Learning dictionaries for information extraction by multi-level bootstrapping
688 Citations1999Ellen Riloff, Rosie Jones
A multilevel bootstrapping algorithm is presented that generates both the semantic lexicon and extraction patterns simultaneously simultaneously and produces high-quality dictionaries for several semantic categories.
Learning Subjective Adjectives from Corpora
520 Citations2000Janyce Wiebe
This paper identifies strong clues of subjectivity using the results of a method for clustering words according to distributional similarity (Lin 1998), seeded by a small amount of detailed manual annotation.
Development and use of a gold-standard data set for subjectivity classifications
498 Citations1999Janyce Wiebe, Rebecca Bruce +1 more
Bias-corrected tags are formulated and successfully used to guide a revision of the coding manual and develop an automatic classifier.
Automatic detection of text genre
358 Citations1997Brett Kessler, Geoffrey Numberg +1 more
A theory of genres as bundles of facets, which correlate with various surface cues, are proposed, and it is argued that genre detection based on surface cues is as successful as Detection based on deeper structural properties.
Journal of PragmaticsGenerating natural language under pragmatic constraints
340 Citations1987Eduard Hovy
arXiv (Cornell University)Tracking point of view in narrative
307 Citations1994Janyce Wiebe
This paper presents an algorithm to develop an algorithm that tracks point of view on the basis of the regularities found in naturally occurring narrative, and describes the results of some preliminary empirical studies, which lend support to the algorithm.
Automatic identification of non-compositional phrases
208 Citations1999Dekang Lin
This work presents a method for automatic identification of non-compositional expressions using their statistical properties in a text corpus based on the hypothesis that when a phrase is non-Compositional, its mutual information differs significantly from the mutual informations of phrases obtained by substituting one of the word in the phrase with a similar word.
arXiv (Cornell University)Automatic Detection of Text Genre
206 Citations1997Brett Kessler, Geoffrey Nunberg +1 more
The Fictions of Language and the Languages of Fiction
176 Citations2003Monika Fludernik
Natural Language EngineeringRecognizing subjectivity: a case study in manual tagging
174 Citations1999Rebecca Bruce, Janyce Wiebe
A case study of a sentence-level categorization in which tagging instructions are developed and used by four judges to classify clauses from the Wall Street Journal as either subjective or objective.
Combining and standardizing large- scale, practical ontologies for machine tranlation and other uses.
166 Citations1998Eduard Hovy
This paper outlines recent attempt at USC/ISI to create a single large Ontology for general free use over the Web, as performed under the aegis of the ANSI Ad Hoc Committee on Ontology Standardization.
A freely available wide coverage morphological analyzer for English
94 Citations1992Daniel Karp, Yves Schabes +2 more
This paper presents a morphological lexicon for English that handle more than 317000 inflected forms derived from over 90000 stems and is the only available free English morphological analyzer with very wide coverage.
Co-occurrence patterns among collocations: a tool for corpus-based lexical knowledge acquisition
93 Citations1993Douglas Biber
It is shown that this method provides information not obtainable through other approaches, including provision of several major senses for each word, an indication of the relationship between collocational patterns, & a more detailed analysis of the senses themselves.
Columbia Academic Commons (Columbia University)The Rules Behind Roles: Identifying speaker role in radio broadcasts
92 Citations2000Regina Barzilay, Michael J. Collins +2 more
An algorithm is implemented that classies story segments into three Speaker Roles based on several content and duration features and correctly classies about 80% of segments when applied to ASR derived transcriptions of broadcast data.
What's yours and what's mine
42 Citations2000Simone Teufel, Marc Moens
The algorithm and a systematic evaluation of a system which can recognize the most salient textual properties that contribute to the global argumentative structure of a text are presented.
Building task-specific interfaces to high volume conversational data
28 Citations1997Loren Terveen, William C. Hill +3 more
A report is reported here on the Phoaks resource recommendation interface, the architecture, and the issues and experience that make up its rationale.
REPRESENTING AND RECOGNIZING POINT OF VIEW
7 Citations1995Warren Sack
The proposed actor-role representation of ideological point of view accords with some recent work by Lakoff (1991), generalizes and improves upon previous AI work on representation of ideology.
National Conference on Artificial IntelligenceAssentor®: An NLP-Based Solution to E-mail Monitoring
6 Citations2000Chinatsu Aone, Mila Ramos-Santacruz +1 more
A quantitative evaluation of applying pattern matching vs. keyword-based searching to e-mail monitoring shows that pattern matching performs significantly better than keyword- based searching both in terms of recall and precision.
