login

Hateful Symbols or Hateful People? Predictive Features for Hate Speech Detection on Twitter

Published 1 January 2016Open access
Zeerak Waseem, Dirk Hovy
Citations1,655
View PDF

TL;DR

A list of criteria founded in critical race theory is provided, and these are used to annotate a publicly available corpus of more than 16k tweets and present a dictionary based the most indicative words in the data.

Abstract

Hate speech in the form of racist and sexist remarks are a common occurrence on social media. For that reason, many social media services address the problem of identifying hate speech, but the definition of hate speech varies markedly and is largely a manual effort. We provide a list of criteria founded in critical race theory, and use them to annotate a publicly available corpus of more than 16k tweets. We analyze the impact of various extra-linguistic features in conjunction with character n-grams for hate-speech detection. We also present a dictionary based the most indicative words in our data.

Keywords

Computer ScienceSocial Sciences