login

Model selection for support vector machine classification

NeurocomputingPublished 12 May 2003Open access
Carl Gold, Peter Sollich
Citations201
View PDF

TL;DR

The results for the evidence gradient ascent method show that also the exact evidence exhibits local optima, but these give test errors which are much less variable and also consistently lower than for the simpler model selection criteria.

Abstract

We address the problem of model selection for Support Vector Machine (SVM)\nclassification. For fixed functional form of the kernel, model selection\namounts to tuning kernel parameters and the slack penalty coefficient $C$. We\nbegin by reviewing a recently developed probabilistic framework for SVM\nclassification. An extension to the case of SVMs with quadratic slack penalties\nis given and a simple approximation for the evidence is derived, which can be\nused as a criterion for model selection. We also derive the exact gradients of\nthe evidence in terms of posterior averages and describe how they can be\nestimated numerically using Hybrid Monte Carlo techniques. Though\ncomputationally demanding, the resulting gradient ascent algorithm is a useful\nbaseline tool for probabilistic SVM model selection, since it can locate maxima\nof the exact (unapproximated) evidence. We then perform extensive experiments\non several benchmark data sets. The aim of these experiments is to compare the\nperformance of probabilistic model selection criteria with alternatives based\non estimates of the test error, namely the so-called ``span estimate'' and\nWahba's Generalized Approximate Cross-Validation (GACV) error. We find that all\nthe ``simple'' model criteria (Laplace evidence approximations, and the Span\nand GACV error estimates) exhibit multiple local optima with respect to the\nhyperparameters. While some of these give performance that is competitive with\nresults from other approaches in the literature, a significant fraction lead to\nrather higher test errors. The results for the evidence gradient ascent method\nshow that also the exact evidence exhibits local optima, but these give test\nerrors which are much less variable and also consistently lower than for the\nsimpler model selection criteria.\n

Keywords

Computer ScienceMathematics