Part-of-Speech Tagging for Twitter: Annotation, Features, and Experiments
FigsharePublished 29 June 2018Open access
Kevin Gimpel, Nathan Schneider, Brendan O’Connor, Dipanjan Das, Daniel P. Mills, Jacob Eisenstein
Citations830
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
Abstract
We address the problem of part-of-speech tagging for English data from the popular microblogging service Twitter. We develop a tagset, annotate data, develop features, and report tagging results nearing 90% accuracy. The data and tools have been made available to the research community with the goal of enabling richer text analysis of Twitter and related social media data sets.
Keywords
Computer Science
ScholarlyCommons (University of Pennsylvania)Conditional Random Fields: Probabilistic Models for Segmenting and Labeling Sequence Data
12,978 Citations2001John Lafferty, Andrew McCallum +1 more
This work presents iterative parameter estimation algorithms for conditional random fields and compares the performance of the resulting models to HMMs and MEMMs on synthetic and natural-language data.
Language<b>WordNet: An electronic lexical database</b> . Ed. by Christiane Fellbaum. Cambridge, MA: MIT Press, 1998. Pp. xxii, 423.
11,687 Citations2000Adam Kilgarriff
The lexical database: nouns in WordNet, George A. Miller modifiers in WordNet, Katherine J. Miller a semantic network of English verbs, and applications of WordNet: building semantic concordances are presented.
Building a Large Annotated Corpus of English: The Penn Treebank
7,528 Citations1993Mitchell P. Marcus
Feature-rich part-of-speech tagging with a cyclic dependency network
2,851 Citations2003Kristina Toutanova, Dan Klein +2 more
A new part-of-speech tagger is presented that demonstrates the following ideas: explicit use of both preceding and following tag contexts via a dependency network representation, broad use of lexical features, and effective use of priors in conditional loglinear models.
Predicting the Future with Social Media
2,064 Citations2010Sitaram Asur, Bernardo A. Huberman
Proceedings of the International AAAI Conference on Web and Social MediaFrom Tweets to Polls: Linking Text Sentiment to Public Opinion Time Series
1,955 Citations2010Brendan O’Connor, Ramnath Balasubramanyan +2 more
This work connects measures of public opinion measured from polls with sentiment measured from text, and finds that temporal smoothing is a critically important issue to support a suc- cessful model.
Word Representations: A Simple and General Method for Semi-Supervised Learning
1,944 Citations2010Joseph Turian, Lev-Arie Ratinov +1 more
This work evaluates Brown clusters, Collobert and Weston (2008) embeddings, and HLBL (Mnih & Hinton, 2009) embeds of words on both NER and chunking, and finds that each of the three word representations improves the accuracy of these baselines.
Robust Sentiment Detection on Twitter from Biased and Noisy Data
875 Citations2010Luciano Barbosa, Junlan Feng
This paper proposes an approach to automatically detect sentiments on Twitter messages (tweets) that explores some characteristics of how tweets are written and meta-information of the words that compose these messages and leverages sources of noisy labels as training data.
Journal of the American Society for Information Science and TechnologySentiment in Twitter events
810 Citations2010Mike Thelwall, Kevan Buckley +1 more
A study of a month of English Twitter posts is reported, assessing whether popular events are typically associated with increases in sentiment strength, as seems intuitively likely and using the top 30 events as a measure of relative increase in (general) term usage.
National Research Council Canada (Government of Canada)Unsupervised Modeling of Twitter Conversations
429 Citations2010Alan Ritter, Colin Cherry +1 more
This work proposes the first unsupervised approach to the problem of modeling dialogue acts in an open domain, trained on a corpus of noisy Twitter conversations, and addresses the challenge of evaluating the emergent model with a qualitative visualization and an intrinsic conversation ordering task.
Proceedings of the International AAAI Conference on Web and Social MediaTweetMotif: Exploratory Search and Topic Summarization for Twitter
370 Citations2010Brendan O’Connor, Michel Krieger +1 more
TweetMotif groups messages by frequent signif- icant terms — a result set’s subtopics — which facili- tate navigation and drilldown through a faceted search interface.
Annotating named entities in Twitter data with crowdsourcing
285 Citations2010Tim Finin, William Murnane +4 more
This paper describes the experience using both Amazon Mechanical Turk (MTurk) and Crowd-Flower to collect simple named entity annotations for Twitter status updates, and describes how to use MTurk to collect judgements on the quality of "word clouds."
Summarizing Microblogs Automatically
170 Citations2010Beaux Sharifi, Mark-Anthony Hutton +1 more
An algorithm is developed that takes a trending phrase or any phrase specified by a user, collects a large number of posts containing the phrase, and provides an automatically created summary of the posts related to the term.
