login

Mistake-Driven Learning in Text Categorization

arXiv (Cornell University)Published 9 June 1997Open access
Ido Dagan, Yael Karov, Dan Roth
Citations33
View PDF

Abstract

Learning problems in the text processing domain often map the text to a space whose dimensions are the measured fea- tures of the text, e.g., its words. Three characteristic properties of this domain are (a) very high dimensionality, (b) both the learned concepts and the instances reside very sparsely in the feature space, and (c) a high variation in the number of active features in an instance. In this work we study three mistake-driven learning algo- rithms for a typical task of this nature - text categorization. We argue

Keywords

Computer Science