Redundancy reduction with information-preserving nonlinear maps
Generate an AI Snapshot to get a quick, structured summary of this paper.
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
This paper generalizes linear principal component analysis to the nonlinear case and implements Barlow's redundancy-reduction principle for unsupervised feature extraction, and drops the usual restriction to Gaussian distributions.
Abstract
The basic idea of linear Principal Component Analyses (PCA) consists in decorrelating coordinates by an orthogonal linear transformation. In this paper we generalize this idea to the nonlinear case. Simultaneously we will drop the usual restriction to gaussian distributions. The linearityand orthogonality condition of linear PCA is substituted with the condition of volume conservation in order to avoid spurious information generated by the nonlinear transformation. This leads us to a still very general class of nonlinear transformations, called symplectic maps. Further on, instead of minimizing the correlation, we minimize the redundancy measured at the output coordinates. This generalizes second order statistics being only valid for gaussian output distributions to higher order statistics. The proposed paradigm implements Barlow's redundancy reduction principle for unsupervised feature extraction. The resulting factorial representation of the joint probability distribution presumably facilitates density estimation and is especially applied to novelty detection.
