Deep Learning via Semi-supervised Embedding
Lecture notes in computer sciencePublished 1 January 2012Open access
Jason Weston, Frédéric Ratle, Hossein Mobahi, Ronan Collobert
Citations568
SJR quartileQ2
SJR score0.35
SNIP0.55
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
Abstract
We show how nonlinear embedding algorithms popular for use with "shallow" semi-supervised learning techniques such as kernel methods can be easily applied to deep multi-layer architectures, either as a regularizer at the output layer, or on each layer of the architecture. This trick provides a simple alternative to existing approaches to deep learning whilst yielding competitive error rates compared to those methods, and existing shallow semi-supervised techniques.
Keywords
Computer ScienceEngineering
Proceedings of the IEEEGradient-based learning applied to document recognition
58,219 Citations1998Yann LeCun, Léon Bottou +2 more
This paper reviews various methods applied to handwritten character recognition and compares them on a standard handwritten digit recognition task, and Convolutional neural networks are shown to outperform all other techniques.
TechnometricsStatistical Learning Theory
26,913 Citations1999Yuhai Wu, Vladimir Vapnik
Presenting a method for determining the necessary and sufficient conditions for consistency of learning process, the author covers function estimates from small data pools, applying these estimations to real-life problems, and much more.
Neural ComputationA Fast Learning Algorithm for Deep Belief Nets
16,404 Citations2006Geoffrey E. Hinton, Simon Osindero +1 more
A fast, greedy algorithm is derived that can learn deep, directed belief networks one layer at a time, provided the top two layers form an undirected associative memory.
ScienceA Global Geometric Framework for Nonlinear Dimensionality Reduction
13,740 Citations2000Joshua B. Tenenbaum, Vin de Silva +1 more
An approach to solving dimensionality reduction problems that uses easily measured local metric information to learn the underlying global geometry of a data set and efficiently computes a globally optimal solution, and is guaranteed to converge asymptotically to the true structure.
Neural ComputationLaplacian Eigenmaps for Dimensionality Reduction and Data Representation
7,709 Citations2003Mikhail Belkin, Partha Niyogi
This work proposes a geometrically motivated algorithm for representing the high-dimensional data that provides a computationally efficient approach to nonlinear dimensionality reduction that has locality-preserving properties and a natural connection to clustering.
PsychometrikaMultidimensional Scaling by Optimizing Goodness of Fit to a Nonmetric Hypothesis
7,398 Citations1964Joseph B. Kruskal
The fundamental hypothesis is that dissimilarities and distances are monotonically related, and a quantitative, intuitively satisfying measure of goodness of fit is defined to this hypothesis.
Machine LearningMultitask Learning
6,312 Citations1997Rich Caruana
Prior work on MTL is reviewed, new evidence that MTL in backprop nets discovers task relatedness without the need of supervisory signals is presented, and new results for MTL with k-nearest neighbor and kernel regression are presented.
The MIT Press eBooksGreedy Layer-Wise Training of Deep Networks
4,704 Citations2007Yoshua Bengio, Pascal Lamblin +2 more
These experiments confirm the hypothesis that the greedy layer-wise unsupervised training strategy mostly helps the optimization, by initializing weights in a region near a good local minimum, giving rise to internal distributed representations that are high-level abstractions of the input, bringing better generalization.
The MIT Press eBooksSemi-Supervised Learning
4,308 Citations2006Olivier Chapelle, Schölkopf, B. +1 more
This first comprehensive overview of semi-supervised learning presents state-of-the-art algorithms, a taxonomy of the field, selected applications, benchmark experiments, and perspectives on ongoing and future research.
Series in machine perception and artificial intelligenceSIGNATURE VERIFICATION USING A “SIAMESE” TIME DELAY NEURAL NETWORK
2,202 Citations1994Jane Bromley, JAMES W. BENTZ +6 more
Unsupervised Learning of Invariant Feature Hierarchies with Applications to Object Recognition
1,104 Citations2007Marc’Aurelio Ranzato, Fu Jie Huang +2 more
An unsupervised method for learning a hierarchy of sparse feature detectors that are invariant to small shifts and distortions that alleviates the over-parameterization problems that plague purely supervised learning procedures, and yields good performance with very few labeled training samples.
Beyond the point cloud
443 Citations2005Vikas Sindhwani, Partha Niyogi +1 more
This paper constructs a family of data-dependent norms on Reproducing Kernel Hilbert Spaces (RKHS) that allow the structure of the RKHS to reflect the underlying geometry of the data.
Deep learning via semi-supervised embedding
387 Citations2008Jason Weston, Frédéric Ratle +1 more
It is shown how nonlinear embedding algorithms popular for use with shallow semi-supervised learning techniques such as kernel methods can be applied to deep multilayer architectures, either as a regularizer at the output layer, or on each layer of the architecture.
Deep learning from temporal coherence in video
355 Citations2009Hossein Mobahi, Ronan Collobert +1 more
A learning method for deep architectures that takes advantage of sequential data, in particular from the temporal coherence that naturally exists in unlabeled video recordings, and is used to improve the performance on a supervised task of interest.
Neural ComputationLearning Optimized Features for Hierarchical Models of Invariant Object Recognition
168 Citations2003Heiko Wersing, Edgar Körner
This work proposes a feedforward model for recognition that shares components like weight sharing, pooling stages, and competitive nonlinearities with earlier approaches but focuses on new methods for learning optimal feature-detecting cells in intermediate stages of the hierarchical network.
View-based 3D object recognition with support vector machines
100 Citations2003Danny Roobaert, Marc M. Van Hulle
This work reports high correct classification of unseen views, especially considering that no domain knowledge is including into the proposed system, and suggests an active learning algorithm to reduce further the required number of training views.
Large scale manifold transduction
45 Citations2008Michael Karlen, Jason Weston +2 more
This work shows how the regularizer of Transductive Support Vector Machines can be trained by stochastic gradient descent for linear models and multi-layer architectures, and proposes a natural generalization of the TSVM loss function that takes into account neighborhood and manifold information directly.
