login

Tokenising, Stemming and Stopword Removal on Anti-spam Filtering Domain

Lecture notes in computer sciencePublished 1 January 2006
José R. Méndez, Eva Iglesias, Florentino Fdez‐Riverola, Fernando Díaz, Juan M. Corchado
Citations55
SJR quartileQ2
SJR score0.35
SNIP0.55

Abstract

Junk e-mail detection and filtering can be considered a cost-sensitive classification problem. Nevertheless, preprocessing methods and noise reduction strategies used to enhance the computational efficiency in text classification cannot be so efficient in e-mail filtering. This fact is demonstrated here where a comparative study of the use of stopword removal, stemming and different tokenising schemes is presented. The final goal is to preprocess the training e-mail corpora of several content-based techniques for spam filtering (machine approaches and case-based systems). Soundness conclusions are extracted from the experiments carried out where different scenarios are taken into consideration.

Keywords

Computer Science