login

A pitfall and solution in multi-class feature selection for text classification

Published 1 January 2004
George Forman
Citations125

TL;DR

Solutions inspired by round-robin scheduling are presented that avoid a pitfall of a large class of feature scoring methods that hurts performance even for a relatively uniform text classification task.

Abstract

Information Gain is a well-known and empirically proven method for high-dimensional feature selection. We found that it and other existing methods failed to produce good results on an industrial text classification problem. On investigating the root cause, we find that a large class of feature scoring methods suffers a pitfall: they can be blinded by a surplus of strongly predictive features for some classes, while largely ignoring features needed to discriminate difficult classes. In this paper we demonstrate this pitfall hurts performance even for a relatively uniform text classification task. Based on this understanding, we present solutions inspired by round-robin scheduling that avoid this pitfall, without resorting to costly wrapper methods. Empirical evaluation on 19 datasets shows substantial improvements.

Keywords

Computer ScienceBiochemistry, Genetics and Molecular Biology