login

Unsupervised Learning Chinese Sentiment Lexicon from Massive Microblog Data

Lecture notes in computer sciencePublished 1 January 2012
Shi Feng, Lin Wang, Weili Xu, Daling Wang, Ge Yu
Citations6
SJR quartileQ2
SJR score0.35
SNIP0.55

TL;DR

A novel method based on occurrence probability with emoticons is presented to learn the candidate sentiment words from the massive microblog data and the accuracy of the learned lexicon is further improved by using the whole microblog space as the corpus.

Abstract

Analyzing people's feelings and emotions in social media has become a major concern for both academic researchers and commercial companies. The sentiment lexicon plays a crucial role in the most sentiment analysis applications. However, existing thesaurus based lexicon building methods suffer from the coverage problems when faced with the new words and new meanings in social media. Nowadays, millions of users share their opinions on different aspects of life everyday in microblogs. In this paper, a novel method based on occurrence probability with emoticons is presented to learn the candidate sentiment words from the massive microblog data and the accuracy of the learned lexicon is further improved by using the whole microblog space as the corpus. Extensive experiments were conducted on real world datasets with different topics. The results show that the proposed method is able to extract the emerging words, and learned lexicon outperforms two well-known Chinese lexicons in classifying the sentiments in microblogs.

Keywords

Computer Science