Model selection for support vector machine classification
Generate an AI Snapshot to get a quick, structured summary of this paper.
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
The results for the evidence gradient ascent method show that also the exact evidence exhibits local optima, but these give test errors which are much less variable and also consistently lower than for the simpler model selection criteria.
Abstract
We address the problem of model selection for Support Vector Machine (SVM)\nclassification. For fixed functional form of the kernel, model selection\namounts to tuning kernel parameters and the slack penalty coefficient $C$. We\nbegin by reviewing a recently developed probabilistic framework for SVM\nclassification. An extension to the case of SVMs with quadratic slack penalties\nis given and a simple approximation for the evidence is derived, which can be\nused as a criterion for model selection. We also derive the exact gradients of\nthe evidence in terms of posterior averages and describe how they can be\nestimated numerically using Hybrid Monte Carlo techniques. Though\ncomputationally demanding, the resulting gradient ascent algorithm is a useful\nbaseline tool for probabilistic SVM model selection, since it can locate maxima\nof the exact (unapproximated) evidence. We then perform extensive experiments\non several benchmark data sets. The aim of these experiments is to compare the\nperformance of probabilistic model selection criteria with alternatives based\non estimates of the test error, namely the so-called ``span estimate'' and\nWahba's Generalized Approximate Cross-Validation (GACV) error. We find that all\nthe ``simple'' model criteria (Laplace evidence approximations, and the Span\nand GACV error estimates) exhibit multiple local optima with respect to the\nhyperparameters. While some of these give performance that is competitive with\nresults from other approaches in the literature, a significant fraction lead to\nrather higher test errors. The results for the evidence gradient ascent method\nshow that also the exact evidence exhibits local optima, but these give test\nerrors which are much less variable and also consistently lower than for the\nsimpler model selection criteria.\n
