login

Feature definition in pattern recognition with small sample size

Pattern RecognitionPublished 1 January 1978
Anil K. Jain, Richard C. Dubes
Citations33
SJR quartileQ1
SJR score2.06
SNIP2.67

TL;DR

The feature definition procedure which is proposed involves partitioning a large set of highly correlated features into subsets, or clusters, through hierarchical clustering and reducing the original set of correlated features to a small set of nearly uncorrelated features.

Abstract

The problem of feature definition in the design of a pattern recognition system where the number of available training samples is small but the number of potential features is excessively large has not received adequate attention. Most of the existing feature extraction and feature selection procedures are not feasible due to computational considerations when the number of features exceeds, say, 100, and are not even applicable when the number of features exceeds the number of patterns. The feature definition procedure which we have proposed involves partitioning a large set of highly correlated features into subsets, or clusters, through hierarchical clustering. Almost any feature selection or extraction procedure, including the constrained maximum variance approach introduced here, can then be applied to each subset to obtain a single representative feature. The original set of correlated features is thus reduced to a small set of nearly uncorrelated features. The utility of this procedure has been demonstrated on a speaker-identification data base which consists of 20 subjects, 156 features, and 180 samples.

Keywords

Computer Science