Feature selection via sensitivity analysis of SVM probabilistic outputs
Machine LearningPublished 2 October 2007Open access
Kaiquan Shen, Chong‐Jin Ong, Xiaoping Li, Einar Wilder‐Smith
Citations84
SJR quartileQ1
SJR score1.15
SNIP2.14
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
The proposed feature-selection method, termed Feature-based Sensitivity of Posterior Probabilities (FSPP), evaluates the importance of a specific feature by computing the aggregate value of the absolute difference of the probabilistic outputs of SVM with and without the feature.
Abstract
10.1007/s10994-007-5025-7
Keywords
Computer Science
Machine LearningRandom Forests
126,192 Citations2001Leo Breiman
Internal estimates monitor error, strength, and correlation and these are used to show the response to increasing the number of features used in the forest, and are also applicable to regression.
ACM Transactions on Intelligent Systems and TechnologyLIBSVM
41,340 Citations2011Chih-Chung Chang, Chih‐Jen Lin
Issues such as solving SVM optimization problems theoretical convergence multiclass classification probability estimates and parameter selection are discussed in detail.
The Nature of Statistical Learning Theory
39,279 Citations1995Vladimir Vapnik
Machine LearningSupport-Vector Networks
33,035 Citations1995Corinna Cortes, Vladimir Vapnik
High generalization ability of support-vector networks utilizing polynomial input transformations is demonstrated and the performance of the support- vector network is compared to various classical learning algorithms that all took part in a benchmark study of Optical Character Recognition.
TechnometricsStatistical Learning Theory
26,913 Citations1999Yuhai Wu, Vladimir Vapnik
Presenting a method for determining the necessary and sufficient conditions for consistency of learning process, the author covers function estimates from small data pools, applying these estimations to real-life problems, and much more.
Cambridge University Press eBooksAn Introduction to Support Vector Machines and Other Kernel-based Learning Methods
13,883 Citations2000Nello Cristianini, John Shawe‐Taylor
Proceedings of the 24th international conference on Machine learning
11,727 Citations2007
This is an index to the papers that appear in the Proceedings of the 29th International Conference on Machine Learning (ICML-12).
A training algorithm for optimal margin classifiers
11,594 Citations1992Bernhard E. Boser, Isabelle Guyon +1 more
A training algorithm that maximizes the margin between the training patterns and the decision boundary is presented, applicable to a wide variety of the classification functions, including Perceptrons, polynomials, and Radial Basis Functions.
Machine LearningGene Selection for Cancer Classification using Support Vector Machines
9,839 Citations2002Isabelle Guyon, Jason Weston +2 more
This paper proposes a new method of gene selection utilizing Support Vector Machine methods based on Recursive Feature Elimination (RFE), and demonstrates experimentally that the genes selected yield better classification performance and are biologically relevant to cancer.
Artificial IntelligenceWrappers for feature subset selection
8,925 Citations1997Ron Kohavi, George H. John
The wrapper method searches for an optimal feature subset tailored to a particular algorithm and a domain and compares the wrapper approach to induction without feature subset selection and to Relief, a filter approach tofeature subset selection.
An introduction to variable and feature selection
6,577 Citations2009Isabelle Guyon
International Conference on Neural Information ProcessingAdvances in kernel methods: support vector learning
5,815 Citations1999Bernhard Schölkopf, Christopher J. C. Burges +1 more
Support vector machines for dynamic reconstruction of a chaotic system, Klaus-Robert Muller et al pairwise classification and support vector machines, Ulrich Kressel.
The MIT Press eBooksFast Training of Support Vector Machines Using Sequential Minimal Optimization
5,462 Citations1998John Platt
Probabilistic Outputs for Support vector Machines and Comparisons to Regularized Likelihood Methods
4,863 Citations1999John Platt
The output of a lassi(cid:12)er should be a alibrated posterior probability to enable post-pro essing and a method to train a kernel lassi with a logit link and a regularized maximum likelihood is proposed.
Technical reportsMaking Large-Scale SVM Learning Practical
4,317 Citations2006Thorsten Joachims
This chapter presents algorithmic and computational results developed for SVM light V 2.0, which make large-scale SVM training more practical and give guidelines for the application of SVMs to large domains.
ePrints Soton (University of Southampton)An Introduction to Support Vector Machines
4,120 Citations2000Nello Cristianini, John Shawe‐Taylor
Fisher discriminant analysis with kernels
2,661 Citations2003Mika Sirén, Gunnar Rätsch +3 more
A non-linear classification technique based on Fisher's discriminant which allows the efficient computation of Fisher discriminant in feature space and large scale simulations demonstrate the competitiveness of this approach.
Machine LearningChoosing Multiple Parameters for Support Vector Machines
2,182 Citations2002Olivier Chapelle, Vladimir Vapnik +2 more
The problem of automatically tuning multiple parameters for pattern recognition Support Vector Machines (SVMs) is considered by minimizing some estimates of the generalization error of SVMs using a gradient descent algorithm over the set of parameters.
The MIT Press eBooksAdvances in Large-Margin Classifiers
1,870 Citations2000
This book provides an overview of recent developments in large margin classifiers, examines connections with other methods, and identifies strengths and weaknesses of the method, as well as directions for future research.
Machine LearningSoft Margins for AdaBoost
1,306 Citations2001Gunnar Rätsch, Takashi Onoda +1 more
It is found that ADABOOST asymptotically achieves a hard margin distribution, i.e. the algorithm concentrates its resources on a few hard-to-learn patterns that are interestingly very similar to Support Vectors.
The Annals of StatisticsClassification by pairwise coupling
1,291 Citations1998Trevor Hastie, Robert Tibshirani
A strategy for polychotomous classification that involves estimating class probabilities for each pair of classes, and then coupling the estimates together is discussed, similar to the Bradley-Terry method for paired comparisons.
Aston Publications Explorer (Aston University)Gaussian Processes for Regression
1,142 Citations1995Christopher K. I. Williams, Carl Edward Rasmussen
This paper investigates the use of Gaussian process priors over functions, which permit the predictive Bayesian analysis for fixed values of hyperparameters to be carried out exactly using matrix operations.
Feature Selection via Concave Minimization and Support Vector Machines
993 Citations1998Paul S. Bradley, O. L. Mangasarian
Numerical tests on 6 public data sets show that classi(cid:12)ers trained by the concave minimization approach and those trained by a support vector machine have comparable 10-fold cross-validation correctness.
Feature Selection for SVMs
977 Citations2000Jason Weston, Sayan Mukherjee +4 more
The resulting algorithms are shown to be superior to some standard feature selection algorithms on both toy data and real-life problems of face recognition, pedestrian detection and analyzing DNA microarray data.
Machine LearningA note on Platt’s probabilistic outputs for support vector machines
868 Citations2007Hsuan-Tien Lin, Chih‐Jen Lin +1 more
An improved algorithm that theoretically converges and avoids numerical difficulties is proposed for Platt’s probabilistic outputs for Support Vector Machines.
Variable selection using svm based criteria
634 Citations2003Alain Rakotomamonjy
Neural ComputationBounds on Error Expectation for Support Vector Machines
627 Citations2000Vladimir Vapnik, Olivier Chapelle
It is proved that the value of the span is always smaller (and can be much smaller) than the diameter of the smallest sphere containing the support vectors, used in previous bounds.
Lecture notes in computer scienceWhich Is the Best Multiclass SVM Method? An Empirical Study
574 Citations2005Kai-Bo Duan, S. Sathiya Keerthi
Empirical evidence is given to show that the one-versus-all method using winner-takes-all strategy and the one to one method implemented by max-wins voting are inferior to another one-Versus-one method: one that uses Platt's posterior probabilities together with the pairwise coupling idea of Hastie and Tibshirani.
IEEE Transactions on Neural NetworksEfficient tuning of SVM hyperparameters using radius/margin bound and iterative algorithms
315 Citations2002S. Sathiya Keerthi
The paper discusses implementation issues related to the tuning of the hyperparameters of a support vector machine (SVM) with L/sub 2/ soft margin, for which the radius/margin bound is taken as the index to be minimized, and iterative techniques are employed for computing radius and margin.
Machine LearningCombined SVM-Based Feature Selection and Classification
247 Citations2005Julia Neumann, Christoph Schnörr +1 more
Four novel continuous feature selection approaches directly minimising the classifier performance are presented, including linear and nonlinear Support Vector Machine classifiers.
IEEE Transactions on Neural NetworksBayesian Support Vector Regression Using a Unified Loss Function
138 Citations2004Wei Chu, S.S. Keerthi +1 more
Experimental results on simulated and real-world data sets indicate that the approach works well even on large data sets, and has the advantages of Bayesian methods for model adaptation and error bars of its predictions.
Pattern Recognition LettersFeature selection algorithms for the generation of multiple classifier systems and their application to handwritten word recognition
58 Citations2004Simon Günter, Horst Bunke
New methods for the creation of classifier ensembles based on feature selection algorithms are introduced, and are evaluated and compared to existing approaches in the context of handwritten word recognition, using a hidden Markov model recognizer as basic classifier.
Minimum Bayes Error Feature Selection for Continuous Speech Recognition
58 Citations2000George Saon, M. Padmanabhan
These approaches outperform standard LDA features and show a 10% relative improvement in the word error rate over state-of-the-art cepstral features on a large vocabulary telephony speech recognition task.
Neural ComputationBayesian Trigonometric Support Vector Classifier
31 Citations2003Wei Chu, S. Sathiya Keerthi +1 more
A novel differentiable loss function is proposed, called the trigonometric loss function, which has the desirable characteristic of natural normalization in the likelihood function, and is followed to set up a Bayesian framework for support vector classification.
