The viability of web-derived polarity lexicons
Published 2 June 2010
Leonid Velikovich, Sasha Blair-Goldensohn, Kerry Hannan, Ryan McDonald
Citations211
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
A lexicon derived from English documents is evaluated, both qualitatively and quantitatively, and it is shown that it provides superior performance to previously studied lexicons, including one derived from WordNet.
Abstract
We examine the viability of building large polarity lexicons semi-automatically from the web. We begin by describing a graph propagation framework inspired by previous work on constructing polarity lexicons from lexical
Keywords
Computer Science
Mining and summarizing customer reviews
7,714 Citations2004Minqing Hu, Bing Liu
This research aims to mine and to summarize all the customer reviews of a product, and proposes several novel techniques to perform these tasks.
Thumbs up?
6,987 Citations2002Bo Pang, Lillian Lee +1 more
This work considers the problem of classifying documents not by topic, but by overall sentiment, e.g., determining whether a review is positive or negative, and concludes by examining factors that make the sentiment classification problem more challenging.
Meeting of the Association for Computational LinguisticsThumbs Up or Thumbs Down? Semantic Orientation Applied to Unsupervised Classification of Reviews
3,654 Citations2002Peter Peter, Turney
Recognizing contextual polarity in phrase-level sentiment analysis
3,379 Citations2005Theresa Wilson, Janyce Wiebe +1 more
A new approach to phrase-level sentiment analysis is presented that first determines whether an expression is neutral or polar and then disambiguates the polarity of the polar expressions.
SENTIWORDNET: A Publicly Available Lexical Resource for Opinion Mining
2,489 Citations2006Andrea Esuli, Fabrizio Sebastiani
SENTIWORDNET is a lexical resource in which each WORDNET synset is associated to three numerical scores Obj, Pos and Neg, describing how objective, positive, and negative the terms contained in the synset are.
Learning from Labeled and Unlabeled Data with Label Propagation
1,568 Citations2002Xiongjie Zhu, Zoubin Ghahramani
A simple iterative algorithm, label propagation, to propagate labels through the dataset along high density areas defined by unlabeled data is proposed and its solution is analyzed, and its connection to several other algorithms is analyzed.
Determining the sentiment of opinions
1,485 Citations2004Soo-Min Kim, Eduard Hovy
A system that, given a topic, automatically finds the people who hold opinions about that topic and the sentiment of each opinion and another module for determining word sentiment and another for combining sentiments within a sentence is presented.
Management ScienceYahoo! for Amazon: Sentiment Extraction from Small Talk on the Web
1,442 Citations2007Sanjiv Ranjan Das, Mike Y. Chen
A methodology for extracting small investor sentiment from stock message boards is developed, which comprises different classifier algorithms coupled together by a voting scheme.
Predicting the semantic orientation of adjectives
1,439 Citations1997Vasileios Hatzivassiloglou, Kathleen McKeown
A log-linear regression model uses constraints from conjunctions to predict whether conjoined adjectives are of same or different orientations, achieving 82% accuracy in this task when each conjunction is considered independently.
Learning extraction patterns for subjective expressions
1,001 Citations2003Ellen Riloff, Janyce Wiebe
A bootstrapping process that learns linguistically rich extraction patterns for subjective (opinionated) expressions while maintaining high precision is presented.
Learning Subjective Adjectives from Corpora
520 Citations2000Janyce Wiebe
This paper identifies strong clues of subjectivity using the results of a method for clustering words according to distributional similarity (Lin 1998), seeded by a small amount of detailed manual annotation.
Lecture notes in computer sciencePulse: Mining Customer Opinions from Free Text
444 Citations2005Michael Gamon, Anthony Aue +2 more
A simple but effective technique for clustering sentences, the application of a bootstrapping approach to sentiment classification, and a novel user-interface are described that enables the exploration of large quantities of customer free text.
Learning Multilingual Subjective Language via Cross-Lingual Projections
399 Citations2007Rada Mihalcea, Carmen Banea +1 more
This paper discusses learning multilingual subjective language via cross-lingual projections with a focus on English as a second language.
Building a Sentiment Summarizer for Local Service Reviews
391 Citations2008Sasha Blair-Goldensohn, Kerry Hannan +4 more
This paper presents a system that summarizes the sen- timent of reviews for a local service such as a restaurant or hotel using aspect-based summarization models, where a summary is built by extracting relevant aspects of a service, such as service or value, aggregating the sentiment per aspect, and selecting aspect-relevant text.
Semi-supervised polarity lexicon induction
340 Citations2009Delip Rao, Deepak Ravichandran
The results indicate that label propagation improves significantly over the baseline and other semi-supervised learning methods like Mincuts and Randomized Mincuts for this task.
Multilingual subjectivity analysis using machine translation
287 Citations2008Carmen Banea, Rada Mihalcea +2 more
Through comparative evaluations on two different languages, it is shown that automatic translation is a viable alternative for the construction of resources and tools for subjectivity analysis in a new target language.
Structured Models for Fine-to-Coarse Sentiment Analysis
286 Citations2007Ryan McDonald, Kerry Hannan +3 more
Experiments show that this structured model for jointly classifying the sentiment of text at varying levels of granularity can significantly reduce classification error relative to models trained in isolation.
Web-scale distributional similarity and entity set expansion
278 Citations2009Patrick Pantel, Eric Crestan +3 more
This work applies the learned similarity matrix to the task of automatic set expansion and presents a large empirical study to quantify the effect on expansion performance of corpus size, corpus quality, seed composition and seed size.
Building Lexicon for Sentiment Analysis from Massive Collection of HTML Documents
268 Citations2007Nobuhiro Kaji, Masaru Kitsuregawa
The key idea is to develop the structural clues so that it achieves extremely high precision at the cost of recall, and build lexicon from the extracted polar sentences.
Generating high-coverage semantic orientation lexicons from overtly marked words and a thesaurus
237 Citations2009Saif M. Mohammad, Cody Dunne +1 more
This work proposes a simple approach to generate a high-coverage semantic orientation lexicon, which includes both individual words and multi-word expressions, using only a Roget-like thesaurus and a handful of affixes and has properties that support the Polyanna Hypothesis.
Adapting a polarity lexicon using integer linear programming for domain-specific sentiment classification
186 Citations2009Yejin Choi, Claire Cardie
A novel method based on integer linear programming that can adapt an existing lexicon into a new one to reflect the characteristics of the data more directly, and Experimental results show that the lexicon adaptation technique improves the performance of fine-grained polarity classification.
Computational IntelligenceMULTI‐DOCUMENT SUMMARIZATION OF EVALUATIVE TEXT
161 Citations2012Giuseppe Carenini, Jackie Chi Kit Cheung +1 more
It is concluded that an effective method for summarizing evaluative arguments must effectively synthe-size the two approaches.
Sentiment summarization
155 Citations2009Kevin Lerman, Sasha Blair-Goldensohn +1 more
A large-scale, end-to-end human evaluation of various sentiment summarization models shows that users have a strong preference for summarizers that model sentiment over non-sentiment baselines, but have no broad overall preference between any of the sentiment-based models.
Generating a non-English subjectivity lexicon
55 Citations2009Valentin Jijkoun, Katja Hofmann
A PageRank-like algorithm is used to bootstrap from the translation of the English lexicon and rank the words in the thesaurus by polarity using the network of lexical relations in Wordnet.
Large-scale computation of distributional similarities for queries
10 Citations2009Enrique Alfonseca, Keith Hall +1 more
It is shown empirically that the large-scale, data-driven approach to computing distributional similarity scores for queries is more effective at ranking query alternatives that the computationally more expensive technique of using the results from a web search engine.
