Sentiment Analysis and Opinion Mining
Synthesis lectures on human language technologiesPublished 23 May 2012
Bing Liu
Citations3,241
SJR quartileQ3
SJR score0.12
SNIP0.00
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
This book is a comprehensive introductory and survey text that covers all important topics and the latest developments in the field with over 400 references and is suitable for students, researchers and practitioners who are interested in social media analysis in general and sentiment analysis in particular.
Abstract
Sentiment analysis and opinion mining is the field of study that analyzes people's opinions, sentiments, evaluations, attitudes, and emotions from written language. It is one of the most active resear
Keywords
Computer Science
Proceedings of the IEEEA tutorial on hidden Markov models and selected applications in speech recognition
22,785 Citations1989L. R. Rabiner
The PageRank Citation Ranking : Bringing Order to the Web
12,645 Citations1999Lawrence M. Page, Sergey Brin +2 more
This paper describes PageRank, a mathod for rating Web pages objectively and mechanically, effectively measuring the human interest and attention devoted to them, and shows how to efficiently compute PageRank for large numbers of pages.
Cambridge University Press eBooksIntroduction to Information Retrieval
10,873 Citations2008Christopher D. Manning, Prabhakar Raghavan +1 more
This textbook teaches classical and web information retrieval, including web search and the related areas of text classification and text clustering from basic concepts, making it perfect for introductory courses in information retrieval for advanced undergraduates and graduate students in computer science.
Foundations of statistical natural language processing
9,996 Citations1999Christopher D. Manning, Hinrich Schütze
Journal of the ACMAuthoritative sources in a hyperlinked environment
9,060 Citations1999Jon Kleinberg
This work proposes and test an algorithmic formulation of the notion of authority, based on the relationship between a set of relevant authoritative pages and the set of “hub pages” that join them together in the link structure, and has connections to the eigenvectors of certain matrices associated with the link graph.
Mining and summarizing customer reviews
7,714 Citations2004Minqing Hu, Bing Liu
This research aims to mine and to summarize all the customer reviews of a product, and proposes several novel techniques to perform these tasks.
Thumbs up?
6,987 Citations2002Bo Pang, Lillian Lee +1 more
This work considers the problem of classifying documents not by topic, but by overall sentiment, e.g., determining whether a review is positive or negative, and concludes by examining factors that make the sentiment classification problem more challenging.
now publishers, Inc. eBooksOpinion Mining and Sentiment Analysis
6,745 Citations2008Bo Pang, Lillian Lee
College Composition and CommunicationA Comprehensive Grammar of the English Language
5,535 Citations1987P Beauvais, Randolph Quirk +3 more
TechnometricsPattern Recognition and Machine Learning
4,637 Citations2007Radford M. Neal
This book covers a broad range of topics for regular factorial designs and presents all of the material in very mathematical fashion and will surely become an invaluable resource for researchers and graduate students doing research in the design of factorial experiments.
Probabilistic latent semantic indexing
3,916 Citations1999Thomas Hofmann
Probabilistic Latent Semantic Indexing is a novel approach to automated document indexing which is based on a statistical latent class model for factor analysis of count data.
Recognizing contextual polarity in phrase-level sentiment analysis
3,379 Citations2005Theresa Wilson, Janyce Wiebe +1 more
A new approach to phrase-level sentiment analysis is presented that first determines whether an expression is neutral or polar and then disambiguates the polarity of the polar expressions.
A sentimental education
3,343 Citations2004Bo Pang, Lillian Lee
A novel machine-learning method is proposed that applies text-categorization techniques to just the subjective portions of the document, which greatly facilitates incorporation of cross-sentence contextual constraints.
Computational LinguisticsLexicon-Based Methods for Sentiment Analysis
3,255 Citations2011Maite Taboada, Julian Brooke +3 more
The Semantic Orientation CALculator (SO-CAL) uses dictionaries of words annotated with their semantic orientation (polarity and strength), and incorporates intensification and negation, and is applied to the polarity classification task.
Machine LearningText Classification from Labeled and Unlabeled Documents using EM
2,749 Citations2000Kamal Nigam, Andrew Kachites McCallum +2 more
This paper shows that the accuracy of learned text classifiers can be improved by augmenting a small number of labeled training documents with a large pool of unlabeled documents, and presents two extensions to the algorithm that improve classification accuracy under these conditions.
SENTIWORDNET: A Publicly Available Lexical Resource for Opinion Mining
2,489 Citations2006Andrea Esuli, Fabrizio Sebastiani
SENTIWORDNET is a lexical resource in which each WORDNET synset is associated to three numerical scores Obj, Pos and Neg, describing how objective, positive, and negative the terms contained in the synset are.
arXiv (Cornell University)Semantic Similarity Based on Corpus Statistics and Lexical Taxonomy
2,224 Citations1997Jay J. Jiang, David W. Conrath
Machine LearningBoosTexter: A Boosting-based System for Text Categorization
2,194 Citations2000Robert E. Schapire, Yoram Singer
This work describes in detail an implementation, called BoosTexter, of the new boosting algorithms for text categorization tasks, and presents results comparing the performance of Boos Texter and a number of other text-categorization algorithms on a variety of tasks.
Seeing stars
2,127 Citations2005Bo Pang, Lillian Lee
A meta-algorithm is applied, based on a metric labeling formulation of the rating-inference problem, that alters a given n-ary classifier's output in an explicit attempt to ensure that similar items receive similar labels.
International Journal of Electronic CommerceThe Effect of On-Line Consumer Reviews on Consumer Purchasing Intention: The Moderating Role of Involvement
2,065 Citations2007Do-Hyung Park, Jumin Lee +1 more
The elaboration likelihood model is used to explain how level of involvement with a product moderates these relationships, and the quality of on-line reviews has a positive effect on consumers' purchasing intention and purchasing intention increases as the number of reviews increases.
Biographies, Bollywood, Boom-boxes and Blenders: Domain Adaptation for Sentiment Classification
2,026 Citations2007John Blitzer, Mark Dredze +1 more
This work extends to sentiment classification the recently-proposed structural correspondence learning (SCL) algorithm, reducing the relative error due to adaptation between domains by an average of 30% over the original SCL algorithm and 46% over a supervised baseline.
Proceedings of the International AAAI Conference on Web and Social MediaFrom Tweets to Polls: Linking Text Sentiment to Public Opinion Time Series
1,955 Citations2010Brendan O’Connor, Ramnath Balasubramanyan +2 more
This work connects measures of public opinion measured from polls with sentiment measured from text, and finds that temporal smoothing is a critically important issue to support a suc- cessful model.
Computers and the HumanitiesAnnotating Expressions of Opinions and Emotions in Language
1,770 Citations2005Janyce Wiebe, Theresa Wilson +1 more
The manual annotation process and the results of an inter-annotator agreement study on a 10,000-sentence corpus of articles drawn from the world press are presented.
IEEE Transactions on Knowledge and Data EngineeringThe Google Similarity Distance
1,764 Citations2007Rudi Cilibrasi, Paul Vitányi
A new theory of similarity between words and phrases based on information distance and Kolmogorov complexity is presented, which is applied to construct a method to automatically extract similarity, the Google similarity distance, of Words and phrases from the WWW using Google page counts.
Extracting product features and opinions from reviews
1,720 Citations2005Ana-Maria Popescu, Oren Etzioni
Opine is introduced, an unsupervised information-extraction system which mines reviews in order to build a model of important product features, their evaluation by reviewers, and their relative quality across products.
Opinion observer
1,604 Citations2005Bing Liu, Minqing Hu +1 more
A novel framework for analyzing and comparing consumer opinions of competing products is proposed, and a new technique based on language pattern mining is proposed to extract product features from Pros and Cons in a particular type of reviews.
Domain adaptation with structural correspondence learning
1,562 Citations2006John Blitzer, Ryan McDonald +1 more
This work introduces structural correspondence learning to automatically induce correspondences among features from different domains in order to adapt existing models from a resource-rich source domain to aresource-poor target domain.
Journal of Interactive MarketingExploring the value of online product reviews in forecasting sales: The case of motion pictures
1,531 Citations2007Chrysanthos Dellarocas, Xiaoquan Zhang +1 more
This study shows that the addition of online product review metrics to a benchmark model that includes prerelease marketing, theater availability and professional critic reviews substantially increases its forecasting accuracy; the forecasting accuracy of the best model outperforms that of several previously published models.
Opinion spam and analysis
1,520 Citations2008Nitin Jindal, Bing Liu
It is shown that opinion spam is quite different from Web spam and email spam, and thus requires different detection techniques, and therefore requires some novel techniques to detect them.
Personality and Social Psychology BulletinLying Words: Predicting Deception from Linguistic Styles
1,457 Citations2003Matthew L. Newman, James W. Pennebaker +2 more
Management ScienceYahoo! for Amazon: Sentiment Extraction from Small Talk on the Web
1,442 Citations2007Sanjiv Ranjan Das, Mike Y. Chen
A methodology for extracting small investor sentiment from stock message boards is developed, which comprises different classifier algorithms coupled together by a voting scheme.
Predicting the semantic orientation of adjectives
1,439 Citations1997Vasileios Hatzivassiloglou, Kathleen McKeown
A log-linear regression model uses constraints from conjunctions to predict whether conjoined adjectives are of same or different orientations, achieving 82% accuracy in this task when each conjunction is considered independently.
A holistic lexicon-based approach to opinion mining
1,388 Citations2008Xiaowen Ding, Bing Liu +1 more
This paper proposes a holistic lexicon-based approach to solving the problem of determining the semantic orientations (positive, negative or neutral) of opinions expressed on product features in reviews by exploiting external evidences and linguistic conventions of natural language expressions.
Semi-Supervised Recursive Autoencoders for Predicting Sentiment Distributions
1,198 Citations2011Richard Socher, Jeffrey Pennington +3 more
A novel machine learning framework based on recursive autoencoders for sentence-level prediction of sentiment label distributions that outperform other state-of-the-art approaches on commonly used datasets, without using any pre-defined sentiment lexica or polarity shifting rules.
Choice Reviews OnlineWeb data mining: exploring hyperlinks, contents, and usage data
1,193 Citations2012
Sentiment analysis
1,158 Citations2003Tetsuya Nasukawa, Jeonghee Yi
This paper illustrates a sentiment analysis approach to extract sentiments associated with polarities of positive or negative for specific subjects from a document, instead of classifying the whole document intopositive or negative.
Towards answering opinion questions
1,051 Citations2003Hong Yu, Vasileios Hatzivassiloglou
A Bayesian classifier for discriminating between documents with a preponderance of opinions such as editorials from regular news stories is presented, and three unsupervised, statistical techniques for the significantly harder task of detecting opinions at the sentence level are described.
Learning extraction patterns for subjective expressions
1,001 Citations2003Ellen Riloff, Janyce Wiebe
A bootstrapping process that learns linguistically rich extraction patterns for subjective (opinionated) expressions while maintaining high precision is presented.
Joint sentiment/topic model for sentiment analysis
972 Citations2009Chenghua Lin, Yulan He
A novel probabilistic modeling framework based on Latent Dirichlet Allocation (LDA), called joint sentiment/topic model (JST), which detects sentiment and topic simultaneously from text is proposed, which is fully unsupervised.
Research Showcase @ Carnegie Mellon University (Carnegie Mellon University)Learning from Labeled and Unlabeled Data using Graph Mincuts
947 Citations2018Avrim Blum, Shuchi Chawla
An algorithm based on finding minimum cuts in graphs, that uses pairwise relationships among the examples in order to learn from both labeled and unlabeled data is considered.
Movie review mining and summarization
909 Citations2006Zhuang Li, Jing Feng +1 more
A multi-knowledge based approach is proposed, which integrates WordNet, statistical analysis and movie knowledge, and the experimental results show the effectiveness of the proposed approach in movie review mining and summarization.
ACM Transactions on Information SystemsSentiment analysis in multiple languages
894 Citations2008Ahmed Abbasi, Hsinchun Chen +1 more
Stylistic features significantly enhanced performance across all testbeds while EWGA also outperformed other feature selection methods, indicating the utility of these features and techniques for document-level classification of sentiments.
International Journal on Digital LibrariesAutomatic recognition of multi-word terms:. the C-value/NC-value method
818 Citations2000Katerina T. Frantzi, Sophia Ananiadou +1 more
This paper presents a domain-independent method for the automatic extraction of multi-word terms, from machine-readable special language corpora, using C-value/NC-value, which enhances the common statistical measure of frequency of occurrence for term extraction, making it sensitive to a particular type ofMulti- word terms, the nested terms.
Learning surface text patterns for a Question Answering system
813 Citations2001Deepak Ravichandran, Eduard Hovy
This paper has developed a method for learning an optimal set of surface text patterns automatically from a tagged corpus, and calculates the precision of each pattern, and the average precision for each question type.
Topic sentiment mixture
813 Citations2007Qiaozhu Mei, Xu Ling +3 more
The proposed Topic-Sentiment Mixture (TSM) model can reveal the latent topical facets in a Weblog collection, the subtopics in the results of an ad hoc query, and their associated sentiments and could also provide general sentiment models that are applicable to any ad hoc topics.
Modeling online reviews with multi-grain topic models
792 Citations2008Ivan Titov, Ryan McDonald
This paper presents a novel framework for extracting ratable aspects of objects from online user reviews and argues that multi-grain models are more appropriate for this task since standard models tend to produce topics that correspond to global properties of objects rather than aspects of an object that tend to be rated by a user.
Cross-domain sentiment classification via spectral feature alignment
789 Citations2010Sinno Jialin Pan, Xiaochuan Ni +3 more
This work develops a general solution to sentiment classification when the authors do not have any labels in a target domain but have some labeled data in a different domain, regarded as source domain and proposes a spectral feature alignment (SFA) algorithm to align domain-specific words from different domains into unified clusters, with the help of domain-independent words as a bridge.
Aspect and sentiment unification model for online review analysis
772 Citations2011Yohan Jo, Alice Oh
This paper proposes Sentence-LDA and extends it to Aspect and Sentiment Unification Model (ASUM), which incorporates aspect and sentiment together to model sentiments toward different aspects and shows that ASUM outperforms other generative models and comes close to supervised classification methods.
Spotting fake reviewer groups in consumer reviews
747 Citations2012Arjun Mukherjee, Bing Liu +1 more
This paper studies spam detection in the collaborative setting, i.e., to discover fake reviewer groups by using several behavioral models derived from the collusion phenomenon among fake reviewers and relation models based on the relationships among groups, individual reviewers, and products they reviewed to detectfake reviewer groups.
Sentiment analyzer: extracting sentiments about a given topic using natural language processing techniques
720 Citations2004Junbo Yi, Tetsuya Nasukawa +2 more
This work presents sentiment analyzer (SA) that extracts sentiment (or opinion) about a subject from online text documents using natural language processing (NLP) techniques.
Improving machine learning approaches to coreference resolution
709 Citations2001Vincent Ng, Claire Cardie
A noun phrase coreference system that extends the work of Soon et al. (2001) and produces the best results to date on the M UC-6 and MUC-7 coreference resolution data sets --- F-measures of 70.4 and 63.4, respectively.
Sentiment Analysis using Support Vector Machines with Diverse Information Sources
693 Citations2004Tony Mullen, Nigel Collier
Experiments on movie review data from the Internet Movie Database demonstrate that hybrid SVMs which combine unigram-style feature-based SVMs with those based on real-valued favorability measures obtain superior performance, producing the best results yet published using this data.
arXiv (Cornell University)Finding Deceptive Opinion Spam by Any Stretch of the Imagination
688 Citations2011Myle Ott, Yejin Choi +2 more
Effects of adjective orientation and gradability on sentence subjectivity
671 Citations2000Vasileios Hatzivassiloglou, Janyce Wiebe
A novel trainable method that statistically combines two indicators of gradability is presented and evaluated, complementing existing automatic techniques for assigning orientation labels.
Latent aspect rating analysis on review text data
659 Citations2010Hongning Wang, Yue Lu +1 more
Empirical experiments show that the proposed latent rating regression model can effectively solve the problem of LARA, and that the detailed analysis of opinions at the level of topical aspects enabled by the proposed model can support a wide range of application tasks, such as aspect opinion summarization, entity ranking based on aspect ratings, and analysis of reviewers rating behavior.
Proceedings of the 19th ACM international conference on Information and knowledge management
620 Citations2010
This year CIKM has received a record high number of submissions, as can be seen from the following statistics: 1382 abstracts submitted, 945 full papers plus 38 demo papers submitted, and 126 papers accepted for presentation as full papers and an additional 165 accepted for short papers.
Measures of distributional similarity
586 Citations1999Lillian Lee
This work presents an empirical comparison of a broad range of measures; a classification of similarity functions based on the information that they incorporate; and the introduction of a novel function that is superior at evaluating potential proxy distributions.
Automatically assessing review helpfulness
522 Citations2006Soo-Min Kim, Patrick Pantel +2 more
This paper considers the task of automatically assessing review helpfulness, and finds that the most useful features include the length of the review, its unigrams, and its product rating.
Partially Supervised Classification of Text Documents
516 Citations2002Bing Liu, Wee Sun Lee +2 more
This paper shows that the problem of identifying documents from a set of documents of a particular topic or class P and a large set M of mixed documents, and that under appropriate conditions, solutions to the constrained optimization problem will give good solution to the partially supervised classification problem.
Discourse ProcessesOn Lying and Being Lied To: A Linguistic Analysis of Deception in Computer-Mediated Communication
514 Citations2007Jeffrey T. Hancock, Lauren E. Curry +2 more
Investigating changes in both the liar's and the conversational partner's linguistic style across truthful and deceptive dyadic communication in a synchronous text-based setting revealed that motivated liars avoided causal terms when lying, whereas unmotivated liars tended to increase their use of negations.
Co-training for cross-lingual sentiment classification
499 Citations2009Xiaojun Wan
A cotraining approach is proposed to making use of unlabeled Chinese data for cross-lingual sentiment classification, which leverages an available English corpus for Chinese sentiment classification by using the English corpus as training data.
Development and use of a gold-standard data set for subjectivity classifications
498 Citations1999Janyce Wiebe, Rebecca Bruce +1 more
Bias-corrected tags are formulated and successfully used to guide a revision of the coding manual and develop an automatic classifier.
Journal of PragmaticsVerbal irony as implicit display of ironic environment: Distinguishing ironic utterances from nonirony
490 Citations2000Akira Utsumi
Opinion Mining and Sentiment Analysis
480 Citations2011Bing Liu
This chapter focuses on mining opinions which indicate positive or negative sentiments, which are of great importance for businesses and consumers wanting to find public or consumer opinions on their products and services.
Extracting opinions, opinion holders, and topics expressed in online news media text
460 Citations2006Soo-Min Kim, Eduard Hovy
This method uses semantic role labeling as an intermediate step to label an opinion holder and topic using data from FrameNet, and decomposes the task into three phases: identifying an opinion-bearing word, labeling semantic roles related to the word in the sentence, and then finding the holder and the topic of the opinion word among the labeled semantic roles.
Review spam detection
447 Citations2007Nitin Jindal, Bing Liu
This paper makes an attempt to study review spam and spam detection.
Identifying comparative sentences in text documents
445 Citations2006Nitin Jindal, Bing Liu
This paper first categorizes comparative sentences into different types, and then presents a novel integrated pattern discovery and supervised learning approach to identifying comparative sentences from text documents.
Sentiment classification on customer feedback data
437 Citations2004Michael Gamon
It is demonstrated that it is possible to perform automatic sentiment classification in the very noisy domain of customer feedback data by using large feature vectors in combination with feature reduction and the addition of deep linguistic analysis features to a set of surface level word n-gram features contributes consistently to classification accuracy.
Just how mad are you? finding strong and weak opinion clauses
435 Citations2004Theresa Wilson, Janyce Wiebe +1 more
This paper presents the first experimental results classifying the strength of opinions and other types of subjectivity and classifies the subjectivity of deeply nested clauses using new syntactic features developed for opinion recognition.
Singapore Management University Institutional Knowledge (InK) (Singapore Management University)Jointly Modeling Aspects and Opinions with a MaxEnt-LDA Hybrid
414 Citations2010Xin Zhao, Jing Jiang +2 more
This paper proposes a MaxEnt-LDA hybrid model to jointly discover both aspects and aspect-specific opinion words and shows that with a relatively small amount of training data, this model can effectively identify aspect and opinion words simultaneously.
Fully automatic lexicon expansion for domain-oriented sentiment analysis
411 Citations2006Hiroshi Kanayama, Tetsuya Nasukawa
This paper proposes an unsupervised lexicon building method for the detection of polar clauses, which convey positive or negative aspects in a specific domain, and its method is robust for corpora in diverse domains and for the size of the initial lexicon.
Extracting semantic orientations of words using spin model
407 Citations2005Hiroya Takamura, Takashi Inui +1 more
The mean field approximation is used to compute the approximate probability function of the system instead of the intractable actual probability function, and a criterion for parameter selection on the basis of magnetization is proposed.
Learning Multilingual Subjective Language via Cross-Lingual Projections
399 Citations2007Rada Mihalcea, Carmen Banea +1 more
This paper discusses learning multilingual subjective language via cross-lingual projections with a focus on English as a second language.
Low-Quality Product Review Detection in Opinion Summarization
392 Citations2007Jingjing Liu, Yunbo Cao +3 more
Experimental results show that the proposed approach effectively discriminates lowquality reviews from high-quality ones and enhances the task of opinion summarization by detecting and filtering low quality reviews.
Building a Sentiment Summarizer for Local Service Reviews
391 Citations2008Sasha Blair-Goldensohn, Kerry Hannan +4 more
This paper presents a system that summarizes the sen- timent of reviews for a local service such as a restaurant or hotel using aspect-based summarization models, where a summary is built by extracting relevant aspects of a service, such as service or value, aggregating the sentiment per aspect, and selecting aspect-relevant text.
Determining the semantic orientation of terms through gloss classification
388 Citations2005Andrea Esuli, Fabrizio Sebastiani
This paper presents a new method for determining the orientation of subjective terms based on the quantitative analysis of the glosses of such terms given in on-line dictionaries, and on the use of the resulting term representations for semi-supervised term classification.
Can online reviews reveal a product's true quality?
382 Citations2006Nan Hu, Paul A. Pavlou +1 more
An analytical model is derived to explain when the mean can serve as a valid representation of a product's true quality, and its implication on marketing practices is discussed.
Rated aspect summarization of short comments
376 Citations2009Yue Lu, ChengXiang Zhai +1 more
The proposed methods are quite general and can be used to generate rated aspect summary automatically given any collection of short comments each associated with an overall rating.
Identifying sources of opinions with conditional random fields and extraction patterns
366 Citations2005Yejin Choi, Claire Cardie +2 more
This work adopts a hybrid approach that combines Conditional Random Fields (Lafferty et al., 2001) and a variation of AutoSlog (Riloff, 1996a), and shows that the combination of these two methods performs better than either one alone.
Phrase dependency parsing for opinion mining
365 Citations2009Yuanbin Wu, Qi Zhang +2 more
A novel approach for mining opinions from product reviews is presented, where it converts opinion mining task to identify product features, expressions of opinions and relations between them by taking advantage of the observation that a lot of product features are phrases.
Proceedings of the International AAAI Conference on Web and Social MediaICWSM — A Great Catchy Name: Semi-Supervised Recognition of Sarcastic Sentences in Online Product Reviews
352 Citations2010Oren Tsur, Dmitry Davidov +1 more
A novel Semi-supervised Algorithm for Sarcasm Identification that recognizes sarcastic sentences in product reviews and speculate on the motivation for using sarcasm in online communities and social networks is presented.
Finding unusual review patterns using unexpected rules
339 Citations2010Nitin Jindal, Bing Liu +1 more
Using the technique, an Amazon.com review dataset is analyzed and many unexpected rules and rule groups which indicate spam activities are found.
The lie detector
339 Citations2009Rada Mihalcea, Carlo Strapparava
It is shown that automatic classification is a viable technique to distinguish between truth and falsehood as expressed in language and a method for class-based feature analysis is introduced, which sheds some light on the features that are characteristic for deceptive text.
Dependency Tree-based Sentiment Classification using CRFs with Hidden Variables
336 Citations2010Tetsuji Nakagawa, Kentaro Inui +1 more
Experimental results of sentiment classification of Japanese and English subjective sentences using conditional random fields with hidden variables showed that the method performs better than other methods based on bag-of-features.
Learning with compositional semantics as structural inference for subsentential sentiment analysis
333 Citations2008Yejin Choi, Claire Cardie
A novel learning-based approach that incorporates structural inference motivated by compositional semantics into the learning procedure is presented, and it is found that expression-level classification accuracy uniformly decreases as additional, potentially disambiguating, context is considered.
<i>Retracted December 5, 2011:</i> A novel lexicalized HMM-based learning framework for web opinion mining
329 Citations2009Wei Jin, Hung Hay Ho
Designing novel review ranking systems
324 Citations2007Anindya Ghose, Panagiotis G. Ipeirotis
It is shown that subjectivity analysis can give useful clues about the helpfulness of a review and about its impact on sales and the results can have several implications for the market design of online opinion forums.
Modeling and Predicting the Helpfulness of Online Reviews
324 Citations2008Yang Liu, Xiangji Huang +2 more
This paper shows that the helpfulness of a review depends on three important factors: the reviewerpsilas expertise, the writing style of the review, and the timeliness of thereview, and presents a nonlinear regression model for helpfulness prediction.
Seeing stars when there aren't many stars
319 Citations2006Andrew B. Goldberg, Xiaojin Zhu
A graph-based semi-supervised learning algorithm is presented to address the sentiment analysis task of rating inference and achieves significantly better predictive accuracy over other methods that ignore the unlabeled examples during training.
Exploiting social context for review quality prediction
309 Citations2010Yue Lu, Panayiotis Tsaparas +2 more
A generic framework for incorporating social context information by adding regularization constraints to the text-based predictor is proposed and has the advantage that the resulting predictor is usable even when social context is unavailable.
ARSA
309 Citations2007Yang Liu, Xiangji Huang +2 more
ARSA is presented, an autoregressive sentiment-aware model, to utilize the sentiment information captured by S-PLSA for predicting product sales performance and is compared with alternative models that do not take into account the sentiment Information.
Conference of the European Chapter of the Association for Computational LinguisticsDetermining Term Subjectivity and Term Orientation for Opinion Mining
306 Citations2006Andrea Esuli, Fabrizio Sebastiani
The task of deciding whether a given term has a positive connotations, or a negative connotation, or has no subjective connotation at all is confronted, and it is shown that determining subjectivity and orientation is a much harder problem than determining orientation alone.
Recognizing stances in online debates
301 Citations2009Swapna Somasundaran, Janyce Wiebe
This paper presents an unsupervised opinion analysis method for debate-side classification, i.e., recognizing which stance a person is taking in an online debate, and shows that this method is substantially better than challenging baseline methods.
Word sense and subjectivity
296 Citations2006Janyce Wiebe, Rada Mihalcea
Empirical evidence is brought in support of the hypotheses that (1) subjectivity is a property that can be associated with word senses, and (2) word sense disambiguation can directly benefit from subjectivity annotations.
Multilingual subjectivity analysis using machine translation
287 Citations2008Carmen Banea, Rada Mihalcea +2 more
Through comparative evaluations on two different languages, it is shown that automatic translation is a viable alternative for the construction of resources and tools for subjectivity analysis in a new target language.
Structured Models for Fine-to-Coarse Sentiment Analysis
286 Citations2007Ryan McDonald, Kerry Hannan +3 more
Experiments show that this structured model for jointly classifying the sentiment of text at varying levels of granularity can significantly reduce classification error relative to models trained in isolation.
…
