login

Hotel Arabic-Reviews Dataset Construction for Sentiment Analysis Applications

Studies in computational intelligencePublished 17 November 2017
Ashraf Elnagar, Yasmin Khalifa, Anas Einea
Citations159
SJR quartileQ4
SJR score0.19
SNIP0.29

TL;DR

This paper introduces HARD (Hotel Arabic-Reviewsdataset), the largest Book Reviews in Arabic Dataset for subjective sentiment analysis and machine language applications, and implements a polarity lexicon-based sentiment analyzer.

Abstract

Arabic language suffers from the lack of available large datasets for machine learning and sentiment analysis applications. This work adds to the recently reported large dataset BRAD, which is the largest Book Reviews in Arabic Dataset. In this paper, we introduce HARD (Hotel Arabic-Reviews Dataset), the largest Book Reviews in Arabic Dataset for subjective sentiment analysis and machine language applications. HARD comprises of 490587 hotel reviews collected from the Booking.com website. Each record contains the review text in the Arabic language, the reviewer's rating on a scale of 1 to 10 stars, and other attributes about the hotel/reviewer. We make available the full unbalanced dataset as well as a balanced subset. To examine the datasets, we implement six popular classifiers using Modern Standard Arabic (MSA) as well as Dialectal Arabic (DA). We test the sentiment analyzers for polarity and rating classifications. Furthermore, we implement a polarity lexicon-based sentiment analyzer. The findings confirm the effectiveness of the classifiers and the datasets. Our core contribution is to make this benchmark-dataset available and accessible to the research community on Arabic language.

Keywords

Computer Science