login

Cross-Lingual Text Categorization

Lecture notes in computer sciencePublished 1 January 2003
Núria Bel, C. H. A. Koster, Marta Villegas
Citations143
SJR quartileQ2
SJR score0.35
SNIP0.55

TL;DR

Practical and cost-effective solutions for automatic Cross-Lingual Text Categorization are described, both in case a sufficient number of training examples is available for each new language and in the case that for some language no training examples are available.

Abstract

This article deals with the problem of Cross-Lingual Text Categorization (CLTC), which arises when documents in different languages must be classified according to the same classification tree. We describe practical and cost-effective solutions for automatic Cross-Lingual Text Categorization, both in case a sufficient number of training examples is available for each new language and in the case that for some language no training examples are available. Experimental results of the bi-lingual classification of the ILO corpus (with documents in English and Spanish) are obtained using bi-lingual training, terminology translation and profile-based translation.

Keywords

Computer Science