login

Efficient Feature Selection via Analysis of Relevance and Redundancy

Published 1 December 2004
Lei Yu, Huan Liu
Citations1,977
SJR quartileQ1
SJR score2.02
SNIP3.07

TL;DR

It is shown that feature relevance alone is insufficient for efficient feature selection of high-dimensional data, and a new framework is introduced that decouples relevance analysis and redundancy analysis.

Abstract

Feature selection is applied to reduce the number of features in many applications where data has hundreds or thousands of features. Existing feature selection methods mainly focus on finding relevant features. In this paper, we show that feature relevance alone is insufficient for efficient feature selection of high-dimensional data. We define feature redundancy and propose to perform explicit redundancy analysis in feature selection. A new framework is introduced that decouples relevance analysis and redundancy analysis. We develop a correlation-based method for relevance and redundancy analysis, and conduct an empirical study of its efficiency and effectiveness comparing with representative methods.

Keywords

Computer ScienceBiochemistry, Genetics and Molecular Biology