login

Optimally Combining Positive and Negative Features for Text Categorization

Published 1 January 2003
Zhaohui Zheng, Rohini K. Srihari
Citations54

TL;DR

The results show that the proposed approach improves text categorization performance, and is in contrast with the standard local feature selection approaches that either only select the terms most indicative of membership or implicitly but not optimally combine the termsmost indicative of members with non-membership.

Abstract

This paper presents a novel local feature selection approach for text categorization. It constructs a feature set for each category by first selecting a set of terms highly indicative of membership as well as another set of terms highly indicative of non-membership, then unifying the two sets. The size ratio of the two sets was empirically chosen to obtain optimal performance. This is in contrast with the standard local feature selection approaches that either (1) only select the terms most indicative of membership; or (2) implicitly but not optimally combine the terms most indicative of membership with non-membership. The experimental comparison between the proposed approach and standard approaches was conducted on four feature selection metrics: chisquare, correlation coefficient, odds ratio, and GSS coefficient. The results show that the proposed approach improves text categorization performance.

Keywords

Computer Science