login

Co-EM support vector learning

Published 1 January 2004
Ulf Brefeld, Tobias Scheffer
Citations200

TL;DR

This work casts linear classifiers into a probabilistic framework and develops a co-EM version of the Support Vector Machine, which conducts experiments on text classification problems and compares the family of semi-supervised support vector algorithms under different conditions, including violations of the assumptions underlying multi-view learning.

Abstract

Multi-view algorithms, such as co-training and co-EM, utilize unlabeled data when the available attributes can be split into independent and compatible subsets. Co-EM outperforms co-training for many problems, but it requires the underlying learner to estimate class probabilities, and to learn from probabilistically labeled data. Therefore, co-EM has so far only been studied with naive Bayesian learners. We cast linear classifiers into a probabilistic framework and develop a co-EM version of the Support Vector Machine. We conduct experiments on text classification problems and compare the family of semi-supervised support vector algorithms under different conditions, including violations of the assumptions underlying multi-view learning. For some problems, such as course web page classification, we observe the most accurate results reported so far.

Keywords

Computer Science