Developing Corpora for Sentiment Analysis: The Case of Irony and Senti-TUT
IEEE Intelligent SystemsPublished 1 March 2013Open access
Cristina Bosco, Viviana Patti, Andrea Bolioli
Citations176
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
Abstract
Senti-TUT—an ongoing Italian project that investigates sentiment and irony in online political discussions—illustrates how to develop corpora for mining and analyzing opinion and sentiment in social media.
Keywords
PsychologyComputer Science
Proceedings of the International AAAI Conference on Web and Social MediaPredicting Elections with Twitter: What 140 Characters Reveal about Political Sentiment
2,681 Citations2010Andranik Tumasjan, Timm O. Sprenger +2 more
It is found that the mere number of messages mentioning a party reflects the election result, and joint mentions of two parties are in line with real world political ties and coalitions.
Computers and the HumanitiesAnnotating Expressions of Opinions and Emotions in Language
1,770 Citations2005Janyce Wiebe, Theresa Wilson +1 more
The manual annotation process and the results of an inter-annotator agreement study on a 10,000-sentence corpus of articles drawn from the world press are presented.
Computational LinguisticsInter-Coder Agreement for Computational Linguistics
1,604 Citations2008Ron Artstein, Massimo Poesio
It is argued that weighted, alpha-like coefficients, traditionally less used than kappa-like measures in computational linguistics, may be more appropriate for many corpus annotation tasks—but that their use makes the interpretation of the value of the coefficient even harder.
Semi-supervised recognition of sarcastic sentences in Twitter and Amazon
422 Citations2010Dmitry Davidov, Oren Tsur +1 more
This paper experiments with semi-supervised sarcasm identification on two very different data sets: a collection of 5.9 million tweets collected from Twitter, and aCollection of 66000 product reviews from Amazon.
Data & Knowledge EngineeringFrom humor recognition to irony detection: The figurative language of social media
371 Citations2012Antonio Reyes, Paolo Rosso +1 more
The research described in this paper is focused on analyzing two playful domains of language: humor and irony, in order to identify key values components for their automatic processing in social media, such as ''tweets''.
Irony in Language and Thought
317 Citations2007Raymond W. Gibbs
Lecture notes in computer scienceThe Hourglass of Emotions
302 Citations2012Erik Cambria, Andrew Livingstone +1 more
A novel biologically-inspired and psychologically-motivated emotion categorisation model that represents affective states both through labels and through four independent but concomitant affective dimensions, which can potentially describe the full range of emotional experiences that are rooted in any of us.
Biomedical Informatics InsightsSentiment Analysis of Suicide Notes: A Shared Task
213 Citations2012John Pestian, Paweł Matykiewicz +7 more
A shared task involving the assignment of emotions to suicide notes resulted in the corpus of fully anonymized clinical text and annotated suicide notes, suggesting that human-like performance on this task is within the reach of currently available technologies.
EmpaTweet: Annotating and Detecting Emotions on Twitter
213 Citations2012Kirk Roberts, Michael Roach +3 more
A corpus collected from Twitter with annotated micro-blog posts annotated at the tweet-level with seven emotions: ANGER, DISGUST, FEAR, JOY, LOVE, SADNESS, and SURPRISE is introduced.
ScopusSenticNet 2: A Semantic and Affective Resource for Opinion Mining and Sentiment Analysis
210 Citations2012Erik Cambria, Catherine Havasi +1 more
By providing the semantics and sentics associated with over 14,000 concepts, SenticNet 2 represents one of the most comprehensive semantic resources for the development of affect-sensitive applications in fields such as social data mining, multimodal affective HCI, and social media marketing.
Irony and Sarcasm: Corpus Generation and Analysis Using Crowdsourcing
180 Citations2012Elena Filatova
A corpus generation experiment is described where regular and sarcastic Amazon product reviews are collected and the resulting corpus can be used for identifying sarcasm on two levels: a document and a text utterance.
arXiv (Cornell University)Tracking Sentiment in Mail: How Genders Differ on Emotional Axes
172 Citations2013Saif M. Mohammad, T.Y. Yang +1 more
Multilingual Subjectivity: Are More Languages Better?
134 Citations2010Carmen Banea, Rada Mihalcea +1 more
This paper explores the integration of features originating from multiple languages into a machine learning approach to subjectivity analysis, and aims to show that this enriched feature set provides for more effective modeling for the source as well as the target languages.
Language Resources and EvaluationPerspectives on crowdsourcing annotations for natural language processing
130 Citations2012Aobo Wang, Cong Duy Vu Hoang +1 more
A faceted analysis of crowdsourcing from a practitioner’s perspective is provided, and how the major crowdsourcing genres fill different parts of this multi-dimensional space is summarized.
Cognitive technologiesEmotion-Oriented Systems
88 Citations2011Roddy Cowie, Catherine Pélachaud +1 more
Computational LinguisticsRelational Features in Fine-Grained Opinion Analysis
68 Citations2012Richard Johansson, Alessandro Moschitti
A set of experiments are described that demonstrate that relational features, mainly derived from dependency-syntactic and semantic role structures, can significantly improve the performance of automatic systems for a number of fine-grained opinion analysis tasks: marking up opinion expressions, finding opinion holders, and determining the polarities of opinion expressions.
Fine-grained German Sentiment Analysis on Social Media
40 Citations2012Saeedeh Momtazi
A fine-grained annotation for German texts is provided, which represents the sentiment strength of the input text using two scores: positive and negative, and a German opinion dictionary of 1,864 words is prepared and compared with other opinion dictionaries for German.
Open Research Online (The Open University)Quantising Opinions for Political Tweets Analysis
24 Citations2012Yulan He, Hassan Saif +2 more
The proposed statistical model for sentiment analysis was able to map the public opinion in Twitter with the actual offline sentiment in real world and show that sentiment analysis based on a simple keyword matching against a sentiment lexicon or a supervised classifier trained with distant supervision does not correlate well with theactual election results.
