Parsimonious Mahalanobis kernel for the classification of high dimensional data
Pattern RecognitionPublished 18 September 2012
Mathieu Fauvel, J. Chanussot, Jón Atli Benediktsson, Alberto Villa
Citations26
SJR quartileQ1
SJR score2.06
SNIP2.67
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
Experimental results show that the proposed kernel based on the Mahalanobis distance is suitable for classifying high dimensional data, providing better classification accuracies than the conventional Gaussian kernel.
Abstract
International audience
Keywords
Computer ScienceEngineering
The Nature of Statistical Learning Theory
39,279 Citations1995Vladimir Vapnik
Data Mining and Knowledge DiscoveryA Tutorial on Support Vector Machines for Pattern Recognition
16,433 Citations1998Christopher J. C. Burges
There are several arguments which support the observed high accuracy of SVMs, which are reviewed and numerous examples and proofs of most of the key theorems are given.
Cambridge University Press eBooksAn Introduction to Support Vector Machines and Other Kernel-based Learning Methods
13,883 Citations2000Nello Cristianini, John Shawe‐Taylor
Multivariate Behavioral ResearchThe Scree Test For The Number Of Factors
13,510 Citations1966Raymond B. Cattell
The Nature of Statistical Learning Theory
10,566 Citations2000Vladimir Vapnik
Journal of Machine Learning ResearchAn introduction to variable and feature selection
7,831 Citations2003GuyonIsabelle, ElisseeffAndré
An introduction to variable and feature selection
6,577 Citations2009Isabelle Guyon
Journal of the Royal Statistical Society Series B (Statistical Methodology)Probabilistic Principal Component Analysis
3,730 Citations1999Michael E. Tipping, Chris Bishop
It is demonstrated how the principal axes of a set of observed data vectors may be determined through maximum likelihood estimation of parameters in a latent variable model that is closely related to factor analysis.
Journal of the Royal Statistical Society Series A (General)Adaptive Control Processes: A Guided Tour.
2,218 Citations1962E. S. Page, Richard Bellman
Machine LearningChoosing Multiple Parameters for Support Vector Machines
2,182 Citations2002Olivier Chapelle, Vladimir Vapnik +2 more
The problem of automatically tuning multiple parameters for pattern recognition Support Vector Machines (SVMs) is considered by minimizing some estimates of the generalization error of SVMs using a gradient descent algorithm over the set of parameters.
IEEE Signal Processing MagazineSpectral unmixing
2,152 Citations2002N. Keshava, John F. Mustard
The outputs of spectral unmixing, endmember, and abundance estimates are important for identifying the material composition of mixtures and the applicability of models and techniques is highly dependent on the variety of circumstances and factors that give rise to mixed pixels.
Lecture notes in computer scienceOn the Surprising Behavior of Distance Metrics in High Dimensional Space
2,058 Citations2001Charų C. Aggarwal, Alexander Hinneburg +1 more
This paper examines the behavior of the commonly used L k norm and shows that the problem of meaningfulness in high dimensionality is sensitive to the value of k, which means that the Manhattan distance metric is consistently more preferable than the Euclidean distance metric for high dimensional data mining applications.
Neural ComputationMixtures of Probabilistic Principal Component Analyzers
1,903 Citations1999Michael E. Tipping, Chris Bishop
PCA is formulated within a maximum likelihood framework, based on a specific form of gaussian latent variable model, which leads to a well-defined mixture model for probabilistic principal component analyzers, whose parameters can be determined using an expectation-maximization algorithm.
The Annals of StatisticsKernel methods in machine learning
1,548 Citations2008Thomas Hofmann, Bernhard Schölkopf +1 more
A review of machine learning methods employing positive definite kernels, ranging from binary classifiers to sophisticated methods for estimation with structured data, which include nonlinear functions as well as functions defined on nonvectorial data.
Information science and statisticsNonlinear Dimensionality Reduction
1,530 Citations2007John A. Lee, Michel Verleysen
The purpose of the book is to summarize clear facts and ideas about well-known methods as well as recent developments in the topic of nonlinear dimensionality reduction, which encompasses many of the recently developed methods.
KERNEL METHODS IN MACHINE LEARNING
1,406 Citations2008Thomas Hofmann, Bernhard Schölkopf +1 more
ACM SIGKDD Explorations NewsletterSubspace clustering for high dimensional data
1,342 Citations2004Lance Parsons, Ehtesham Haque +1 more
A survey of the various subspace clustering algorithms along with a hierarchy organizing the algorithms by their defining characteristics is presented, comparing the two main approaches using empirical scalability and accuracy tests and discussing some potential applications where sub space clustering could be particularly useful.
ACM Transactions on Knowledge Discovery from DataClustering high-dimensional data
1,088 Citations2009Hans‐Peter Kriegel, Peer Kröger +1 more
This survey tries to clarify the different problem definitions related to subspace clustering in general; the specific difficulties encountered in this field of research; the varying assumptions, heuristics, and intuitions forming the basis of different approaches; and how several prominent solutions tackle different problems.
EconometricaAdaptive Control Processes: A Guided Tour
1,020 Citations1962Dimitri Morgenstern, Richard Bellman
Pattern RecognitionA survey of kernel and spectral methods for clustering
818 Citations2007Maurizio Filippone, Francesco Camastra +2 more
A survey of kernel and spectral clustering methods, two approaches able to produce nonlinear separating hypersurfaces between clusters and an explicit proof of the fact that these two paradigms have the same objective is reported.
Signal Theory Methods in Multispectral Remote Sensing
818 Citations2003D. A. Landgrebe
This work focuses on the development of a Models for Multispectral Image Data Preprocessing, which combines Hyperspectral Data Characteristics with Probability Theory.
IEEE Transactions on Systems Man and Cybernetics Part C (Applications and Reviews)Supervised classification in high-dimensional space: geometrical, statistical, and asymptotical properties of multivariate data
419 Citations1998Luís Ortiz Jiménez, D. A. Landgrebe
High-dimensional space properties are investigated using Euclidean and Cartesian geometry and their implication for high-dimensional data and its analysis is studied in order to illuminate the differences between conventional spaces and hyperdimensional space.
Computational Statistics & Data AnalysisHigh-dimensional data clustering
367 Citations2007Charles Bouveyron, Stéphane Girard +1 more
A family of Gaussian mixture models designed for high-dimensional data which combine the ideas of dimension reduction and parsimonious modeling give rise to a clustering method based on the Expectation-Maximization algorithm which is called High-Dimensional Data Clustering (HDDC).
IEEE Transactions on Neural NetworksEfficient tuning of SVM hyperparameters using radius/margin bound and iterative algorithms
315 Citations2002S. Sathiya Keerthi
The paper discusses implementation issues related to the tuning of the hyperparameters of a support vector machine (SVM) with L/sub 2/ soft margin, for which the radius/margin bound is taken as the index to be minimized, and iterative techniques are employed for computing radius and margin.
IEEE Transactions on Knowledge and Data EngineeringThe Concentration of Fractional Distances
310 Citations2007Damien François, Vincent Wertz +1 more
This paper justifies the use of alternative distances to fight concentration by showing that the concentration is indeed an intrinsic property of the distances and not an artifact from a finite sample, and an estimation of the concentration as a function of the exponent of the distance and of the distribution of the data.
Pattern RecognitionA spatial–spectral kernel-based approach for the classification of remote-sensing images
278 Citations2011Mathieu Fauvel, J. Chanussot +1 more
The proposed method deals with the joint use of the spatial and the spectral information provided by the remote-sensing images with very high spatial resolution and is competitive with other contextual methods.
Foundations and Trends® in Machine LearningDimension Reduction: A Guided Tour
276 Citations2010Christopher J. C. Burges
A tutorial overview of several geometric methods for dimension reduction by dividing the methods into projective methods and methods that model the manifold on which the data lies.
The Curse of Highly Variable Functions for Local Kernel Machines
193 Citations2005Yoshua Bengio, Olivier Delalleau +1 more
A series of theoretical arguments are presented supporting the claim that a large class of modern learning algorithms that rely solely on the smoothness prior - with similarity between examples expressed with a local kernel - are sensitive to the curse of dimensionality, or more precisely to the variability of the target.
IEEE Transactions on Pattern Analysis and Machine IntelligenceMixtures of Factor Analyzers with Common Factor Loadings: Applications to the Clustering and Visualization of High-Dimensional Data
159 Citations2009Jangsun Baek, Geoffrey J. McLachlan +1 more
Pattern RecognitionA new approach to mixed pixel classification of hyperspectral imagery based on extended morphological profiles
130 Citations2004Antonio Plaza, Pablo Martinez +2 more
A quantitative and comparative performance study with regards to other standard hyperspectral analysis methodologies reveals that the combined utilization of spatial and spectral information in the proposed technique produces classification results which are superior to those found by using the spectral information alone.
Communication in Statistics- Theory and MethodsHigh-Dimensional Discriminant Analysis
117 Citations2007Charles Bouveyron, Stéphane Girard +1 more
A new parameterization of the Gaussian model which combines the ideas of dimension reduction and constraints on the model is proposed, which takes into account the specific subspace and the intrinsic dimension of each class to limit the number of parameters to estimate.
Pattern RecognitionStatistical pattern recognition in remote sensing
104 Citations2008Chi Hau Chen, Pei-Gee Peter Ho
Though the paper is largely tutorial in nature, some specific issues considered are image models for characterization of contextual information, neural networks for image classification, and the performance measures.
Learning high-dimensional data
95 Citations2001Michel Verleysen
Pattern RecognitionSVM-based feature extraction for face recognition
75 Citations2010Sang-Ki Kim, Youn Jung Park +2 more
This paper redesigns the between-class scatter matrix based on the SVM margins to facilitate an effective and reliable feature extraction and follows by a regularization of the within- class scatter matrix.
IEEE Transactions on Neural NetworksWeighted Mahalanobis Distance Kernels for Support Vector Machines
72 Citations2007Defeng Wang, Daniel Yeung +1 more
This paper first finds the data structure for each class adaptively in the input space via agglomerative hierarchical clustering, and then construct the weighted Mahalanobis distance (WMD) kernels using the detected data distribution information.
Lecture notes in computer scienceOn the effects of dimensionality on data analysis with neural networks
57 Citations2003Michel Verleysen, D. François +2 more
Some limitations of conventional concepts in high-dimensional data analysis are shown and some research directions are suggested as the use of alternative distance definitions and of non-linear dimension reduction.
Pattern Recognition LettersIntrinsic dimension estimation by maximum likelihood in isotropic probabilistic PCA
48 Citations2011Charles Bouveyron, Gilles Celeux +1 more
The surprising result of the asymptotic consistency of the maximum likelihood criterion for determining the intrinsic dimension of a dataset in an isotropic version of probabilistic principal component analysis (PPCA) is demonstrated.
Lecture notes in computer scienceTraining of Support Vector Machines with Mahalanobis Kernels
48 Citations2005Shigeo Abe
This paper proposes using Mahalanobis kernels, which are generalized RBF kernels, to solve the problem of optimizing the kernel parameter and the margin parameter by time-consuming cross validation for model selection.
IEEE Transactions on Neural NetworksA Geometrical Method to Improve Performance of the Support Vector Machine
41 Citations2007Peter M. Williams, Sheng Li +2 more
This letter investigates a geometrical method to optimize the kernel function, a modification of the one proposed by S. Amari and S. Wu, that works efficiently and overcomes the susceptibility of the original method.
IEEE Transactions on Pattern Analysis and Machine IntelligenceDistance-preserving projection of high-dimensional data for nonlinear dimensionality reduction
39 Citations2004Yang Li
A distance-preserving method is presented to map high-dimensional data sequentially to low-dimensional space and preserves exact distances of each data point to its nearest neighbor and to some other near neighbors.
Hyperspectral image classification with mahalanobis relevance vector machines
13 Citations2007Gustau Camps‐Valls, Antonio Rodrigo-González +3 more
The Mahalanobis kernel is included in the formulation of the RVM to take into account the covariance of the features in the classification process, and also the ease of free parameters tuning.
Lecture notes in computer scienceSupport Vector Regression Using Mahalanobis Kernels
11 Citations2006Yuya Kamada, Shigeo Abe
This paper determines the covariance matrix for the Mahalanobis kernel using all the training data and estimation performance is comparable to or better than that of an RBF kernel optimized by grid search.
The European Symposium on Artificial Neural NetworksSearching for the embedded manifolds in high- dimensional data, problems and unsolved questions
11 Citations2002Jeanny Hérault, Anne Guérin-Dugué +1 more
A bird's eye view over various techniques of data reduction, from linear multidimensional scaling to non-linear and non-parametric methods is given.
Mahalanobis kernel for the classification of hyperspectral images
8 Citations2010Mathieu Fauvel, Alberto Villa +2 more
Results on real data sets empirically demonstrate that the proposed Mahalanobis kernel leads to an increase of the classification accuracy by comparison to standard kernels.
