Development and use of a gold-standard data set for subjectivity classifications
Published 1 January 1999Open access
Janyce Wiebe, Rebecca Bruce, Thomas P. O'Hara
Citations498
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
Bias-corrected tags are formulated and successfully used to guide a revision of the coding manual and develop an automatic classifier.
Abstract
This paper presents a case study of analyzing and improving intercoder reliability in discourse tagging using statistical techniques. Bias-corrected tags are formulated and successfully used to guide a revision of the coding manual and develop an automatic classifier.
Keywords
Computer Science
Journal of the Royal Statistical Society Series B (Statistical Methodology)Maximum Likelihood from Incomplete Data Via the <i>EM</i> Algorithm
49,657 Citations1977A. P. Dempster, N. M. Laird +1 more
Journal of the American Statistical AssociationContent Analysis: An Introduction to its Methodology.
24,566 Citations1984Mack Shelley, Klaus Krippendorff
Building a Large Annotated Corpus of English: The Penn Treebank
7,528 Citations1993Mitchell P. Marcus
College Composition and CommunicationA Comprehensive Grammar of the English Language
5,535 Citations1987P Beauvais, Randolph Quirk +3 more
Content Analysis: An Introduction to Its Methodology
5,456 Citations2019Klaus Krippendorff
Contemporary Sociology A Journal of ReviewsDiscrete Multivariate Analysis: Theory and Practice.
5,160 Citations1975James R. Beniger, Yvonne Bishop +4 more
Accurate methods for the statistics of surprise and coincidence
2,688 Citations1993Ted Dunning
arXiv (Cornell University)Assessing agreement on classification tasks: the kappa statistic
2,092 Citations1996Jean Carletta
Journal of the Royal Statistical Society Series C (Applied Statistics)Maximum Likelihood Estimation of Observer Error-Rates Using the EM Algorithm
1,666 Citations1979A. P. Dawid, A. M. Skene
The EM algorithm is shown to provide a slow but sure way of obtaining maximum likelihood estimates of the parameters of interest in compiling a patient record.
BiometrikaExploratory latent structure analysis using both identifiable and unidentifiable models
1,585 Citations1974Leo A. Goodman
Predicting the semantic orientation of adjectives
1,439 Citations1997Vasileios Hatzivassiloglou, Kathleen McKeown
A log-linear regression model uses constraints from conjunctions to predict whether conjoined adjectives are of same or different orientations, achieving 82% accuracy in this task when each conjunction is considered independently.
News as Discourse
1,168 Citations1988Teun Adrianus van Dijk
Bayesian classification (AutoClass): theory and results
972 Citations1996Peter Cheeseman, John Stutz
It is emphasized that no current unsupervised classi(cid:12)cation system can produce maximally useful results when operated alone, and that it is the interaction between domain experts and the machine searching over the model space, that generates new knowledge.
Comparative LiteratureUnspeakable Sentences: Narration and Representation in the Language of Fiction
840 Citations1984Gerald Prince, Ann Banfield
Choice Reviews OnlineGoodness-of-fit statistics for discrete multivariate data
744 Citations1989
This goodness of fit statistics for discrete multivariate data helps people to read a good book with a cup of tea in the afternoon, instead they are facing with some infectious virus inside their laptop.
Journal of PragmaticsGenerating natural language under pragmatic constraints
340 Citations1987Eduard Hovy
arXiv (Cornell University)Tracking point of view in narrative
307 Citations1994Janyce Wiebe
This paper presents an algorithm to develop an algorithm that tracks point of view on the basis of the regularities found in naturally occurring narrative, and describes the results of some preliminary empirical studies, which lend support to the algorithm.
Natural Language EngineeringRecognizing subjectivity: a case study in manual tagging
174 Citations1999Rebecca Bruce, Janyce Wiebe
A case study of a sentence-level categorization in which tagging instructions are developed and used by four judges to classify clauses from the Wall Street Journal as either subjective or objective.
ArXiv.orgLinear Segmentation and Segment Significance
101 Citations1998Min‐Yen Kan, Judith L. Klavans +1 more
A new method for discovering a segmental discourse structure of a document while categorizing each segment's function and importance is presented, using a zero-sum weighting scheme.
Dialogue act tagging with Transformation-Based Learning
98 Citations1998Ken Samuel, Sandra Carberry +1 more
This work extracts values of well-motivated features of utterances, such as speaker direction, punctuation marks, and a new feature, called dialogue act cues, which it finds to be more effective than cue phrases and word n-grams in practice.
Journal of Artificial Intelligence ResearchCue Phrase Classification Using Machine Learning
69 Citations1996Diane Litman
This paper explores the use of machine learning for classifying cue phrases as discourse or sentential in natural language processing systems that exploit discourse structure, e.g., for performing tasks such as anaphora resolution and plan recognition.
Knowledge lean word sense disambiguation
60 Citations1997Ted Pedersen
A corpus-based approach to word-sense disambiguation that only requires information that can be automatically extracted from untagged text and Gibbs Sampling results in small but consistent improvement in disambigsuation accuracy over the EM algorithm.
Computational LinguisticsDecomposable modeling in natural language processing
42 Citations1999Rebecca Bruce, Janyce Wiebe
A framework for developing probabilistic classifiers in natural language processing by formulating models that capture the most important interdependencies among features, to avoid overfitting the data while also characterizing the data well is described.
Proceedings of the 17th international conference on Computational linguistics -Dialogue act tagging with Transformation-Based Learning
35 Citations1998Ken Samuel, Sandra Carberry +1 more
Word-Sense Distinguishability and Inter-Coder Agreement
28 Citations1998Rebecca Bruce, Janyce Wiebe
An approach to analyze the sense tags assigned by five judgps to the noun intcr·est for the purpose of formulating a refined and more reliable set of category designations is presented.
Mapping Collocational Properties into Machine Learning Features
15 Citations1998Janyce Wiebe, Kenneth J. McKeever +1 more
A statistical analysis of the results across di(cid:11)erent machine learning algorithms shows the relationship between property and organization was strikingly consistent across algorithms.
arXiv (Cornell University)Probabilistic Event Categorization
7 Citations1997Janyce Wiebe, Rebecca Bruce +1 more
This paper describes the automation of a new text categorization task and presents various types of properties experimented with in this work, including a way to take advantage of properties that are low frequency but strongly indicative of a class.
