login

A Comparative Study on Feature Selection in Chinese Text Categorization

Zhongwen xinxi xuebaoPublished 1 January 2004
Dai Liu
Citations50

TL;DR

A furthermore experiment proved that the combined feature selection method is effective, finding IG, MI and CHI had poor performance in the test, though they behave well in English text categorization.

Abstract

This paper is a comparative study of feature selection methods in text categorization. Four methods were evaluated, including document frequency (DF), information gain (IG), mutual information (MI) and χ 2 test (CHI). A Support Vector Machine ( SVM) and a k nearest neighbor (KNN) were selected as the evaluating classifiers. We found IG, MI and CHI had poor performance in our test, though they behave well in English text categorization. We analyzed the reasons theoretically and put forwarded the possible solutions. A furthermore experiment proved that the combined feature selection method is effective.

Keywords

Computer Science