login

Automatically annotating a five-billion-word corpus of Japanese blogs for sentiment and affect analysis

Computer Speech & LanguagePublished 18 May 2013Open access
Michał Ptaszyński, Rafał Rzepka, Kenji Araki, Yoshio Momouchi
Citations29
View PDF

TL;DR

This paper presents research on automatic annotation of a five-billion-word corpus of Japanese blogs with information on affect and sentiment, which is applied in several tasks, such as generation of emotion object ontology or retrieval of emotional and moral consequences of actions.

Abstract

This paper presents our research on automaticannotation of a five-billion-word corpus ofJapanese blogs with information on affect andsentiment. We first perform a study in emotionblog corpora to discover that there has beenno large scale emotion corpus available forthe Japanese language. We choose the largestblog corpus for the language and annotate itwith the use of two systems for affect analysis:ML-Ask for word- and sentence-levelaffect analysis and CAO for detailed analysisof emoticons. The annotated informationincludes affective features like sentencesubjectivity (emotive/non-emotive) or emotionclasses (joy, sadness, etc.), useful in affectanalysis. The annotations are also generalizedon a 2-dimensional model of affect to obtaininformation on sentence valence/polarity(positive/negative) useful in sentiment analysis.The annotations are evaluated in severalways. Firstly, on a test set of a thousand sentencesextracted randomly and evaluated byover forty respondents. Secondly, the statisticsof annotations are compared to other existingemotion blog corpora. Finally, the corpus isapplied in several tasks, such as generation ofemotion object ontology or retrieval of emotionaland moral consequences of actions.

Keywords

PsychologyComputer Science