login

Neighbor-weighted K-nearest neighbor for unbalanced text corpus

Expert Systems with ApplicationsPublished 13 January 2005
S TAN
Citations344
SJR quartileQ1
SJR score1.85
SNIP2.55

TL;DR

The neighbor-weighted K-nearest neighbor algorithm, i.e. NWKNN, is proposed, which achieves significant classification performance improvement on imbalanced corpora.

Abstract

Text categorization or classification is the automated assigning of text documents to pre-defined classes based on their contents. Many of classification algorithms usually assume that the training examples are evenly distributed among different classes. However, unbalanced data sets often appear in many practical applications. In order to deal with uneven text sets, we propose the neighbor-weighted K-nearest neighbor algorithm, i.e. NWKNN. The experimental results indicate that our algorithm NWKNN achieves significant classification performance improvement on imbalanced corpora.

Keywords

Computer Science