DRAW: A Recurrent Neural Network For Image Generation
arXiv (Cornell University)Published 16 February 2015Open access
Karol Gregor, Ivo Danihelka, Alex Graves, Danilo Jimenez Rezende, Daan Wierstra
Citations963
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
Abstract
This paper introduces the Deep Recurrent Attentive Writer (DRAW) neural network architecture for image generation. DRAW networks combine a novel spatial attention mechanism that mimics the foveation of the human eye, with a sequential variational auto-encoding framework that allows for the iterative construction of complex images. The system substantially improves on the state of the art for generative models on MNIST, and, when trained on the Street View House Numbers dataset, it generates images that cannot be distinguished from real data with the naked eye.
Keywords
Computer Science
Neural ComputationLong Short-Term Memory
98,079 Citations1997Sepp Hochreiter, Jürgen Schmidhuber
A novel, efficient, gradient based method called long short-term memory (LSTM) is introduced, which can learn to bridge minimal time lags in excess of 1000 discrete-time steps by enforcing constant error flow through constant error carousels within special units.
UvA-DARE (University of Amsterdam)Adam: A Method for Stochastic Optimization
84,783 Citations2014Diederik P. Kingma, Jimmy Ba
Proceedings of the IEEEGradient-based learning applied to document recognition
58,219 Citations1998Yann LeCun, Léon Bottou +2 more
This paper reviews various methods applied to handwritten character recognition and compares them on a standard handwritten digit recognition task, and Convolutional neural networks are shown to outperform all other techniques.
ScienceReducing the Dimensionality of Data with Neural Networks
20,953 Citations2006Geoffrey E. Hinton, Ruslan Salakhutdinov
UvA-DARE (University of Amsterdam)Auto-Encoding Variational Bayes
15,628 Citations2013Diederik P. Kingma, Max Welling
A stochastic variational inference and learning algorithm that scales to large datasets and, under some mild differentiability conditions, even works in the intractable case is introduced.
Neural ComputationLearning to Forget: Continual Prediction with LSTM
5,419 Citations2000Felix A. Gers, Jürgen Schmidhuber +1 more
This work identifies a weakness of LSTM networks processing continual input streams that are not a priori segmented into subsequences with explicitly marked ends at which the network's internal state could be reset, and proposes a novel, adaptive forget gate that enables an LSTm cell to learn to reset itself at appropriate times, thus releasing internal resources.
Reading digits in natural images with unsupervised feature learning
4,546 Citations2024Yuval Netzer
A new benchmark dataset for research use is introduced containing over 600,000 labeled digits cropped from Street View images, and variants of two recently proposed unsupervised feature learning methods are employed, finding that they are convincingly superior on benchmarks.
Leibniz-Zentrum für Informatik (Schloss Dagstuhl)LOL: An Investigation into Cybernetic Humor, or: Can Machines Laugh?
3,084 Citations2016Alex Graves, Gervasi, Vincenzo +1 more
This paper shows how Long Short-term Memory recurrent neural networks can be used to generate complex sequences with long-range structure, simply by predicting one data point at a time.
arXiv (Cornell University)Stochastic Backpropagation and Approximate Inference in Deep Generative Models
2,644 Citations2014Danilo Jimenez Rezende, Shakir Mohamed +1 more
arXiv (Cornell University)Recurrent Models of Visual Attention
2,272 Citations2014Volodymyr Mnih, Nicolas Heess +2 more
A novel recurrent neural network model that is capable of extracting information from an image or video by adaptively selecting a sequence of regions or locations and only processing the selected regions at high resolution is presented.
Deep Boltzmann machines
1,778 Citations2009Ruslan Salakhutdinov, Geoffrey E. Hinton
A new learning algorithm for Boltzmann machines that contain many layers of hidden variables that is made more efficient by using a layer-by-layer “pre-training” phase that allows variational inference to be initialized with a single bottomup pass.
Neural ComputationThe Helmholtz Machine
1,219 Citations1995Peter Dayan, Geoffrey E. Hinton +2 more
A way of finessing this combinatorial explosion by maximizing an easily computed lower bound on the probability of the observations is described, viewed as a form of hierarchical self-supervised learning that may relate to the function of bottom-up and top-down cortical processing pathways.
arXiv (Cornell University)Multiple Object Recognition with Visual Attention
701 Citations2014Jimmy Ba, Volodymyr Mnih +1 more
On the quantitative analysis of deep belief networks
432 Citations2008Ruslan Salakhutdinov, Iain Murray
It is shown that Annealed Importance Sampling (AIS) can be used to efficiently estimate the partition function of an RBM, and a novel AIS scheme for comparing RBM's with different architectures is presented.
Edinburgh Research Explorer (University of Edinburgh)The Neural Autoregressive Distribution Estimator
431 Citations2011Hugo Larochelle, Iain Murray
A new approach for modeling the distribution of high-dimensional vectors of discrete variables inspired by the restricted Boltzmann machine, which outperforms other multivariate binary distribution estimators on several datasets and performs similarly to a large (but intractable) RBM.
Learning to combine foveal glimpses with a third-order Boltzmann machine
427 Citations2010Hugo Larochelle, Geoffrey E. Hinton
A model based on a Boltzmann machine with third-order connections that can learn how to accumulate information about a shape over several fixations is described, showing that it can perform at least as well as a model trained on whole images.
arXiv (Cornell University)Multi-digit Number Recognition from Street View Imagery using Deep Convolutional Neural Networks
284 Citations2013Ian Goodfellow, Yaroslav Bulatov +3 more
This paper employs the DistBelief implementation of deep neural networks in order to train large, distributed neural networks on high quality images and finds that the performance of this approach increases with the depth of the convolutional network.
Neural ComputationLearning Where to Attend with Deep Architectures for Image Tracking
200 Citations2012Misha Denil, Loris Bazzani +2 more
An attentional model for simultaneous object tracking and recognition that is driven by gaze data is discussed, and a straightforward extension of the existing approach to the partial information setting results in poor performance, and an alternative method based on modeling the reward surface as a gaussian process is proposed.
arXiv (Cornell University)Attention for Fine-Grained Categorization
119 Citations2014Pierre Sermanet, Andrea Frome +1 more
This paper presents experiments extending the work of Ba et al. (2014) on recurrent neural models for attention into less constrained visual environments, specifically fine-grained categorization on the Stanford Dogs data set using an RNN of the same structure but substitute a more powerful visual network and perform large-scale pre-training of the visual network outside of the attention RNN.
arXiv (Cornell University)Deep AutoRegressive Networks
110 Citations2013Karol Gregor, Ivo Danihelka +3 more
An efficient approximate parameter estimation method based on the minimum description length (MDL) principle is derived, which can be seen as maximising a variational lower bound on the log-likelihood, with a feedforward neural network implementing approximate inference.
arXiv (Cornell University)A Deep and Tractable Density Estimator
103 Citations2013Benigno Uría, Iain Murray +1 more
arXiv (Cornell University)On Learning Where To Look
37 Citations2014Marc’Aurelio Ranzato
This work describes a learning based method that recognizes objects through a series of glimpses that performs an amount of computation that scales with the complexity of the input rather than its number of pixels.
International Journal of Computer VisionA Neural Autoregressive Approach to Attention-based Recognition
33 Citations2014Yin Zheng, Richard S. Zemel +2 more
This paper proposes an alternative approach based on a feed-forward, auto-regressive architecture, which permits exact calculation of training gradients (given the fixation sequence), unlike for the RBM model.
Belarusian State Pedagogical University repository (Belarusian State Pedagogical University)Optimizing Neural Networks that Generate Iimages
29 Citations2014Tijmen Tieleman
This work leverages the ability to generate images, for the purpose of recognizing other images, based on the reverse process: generating images.
UvA-DARE (University of Amsterdam)Markov Chain Monte Carlo and Variational Inference: Bridging the Gap
23 Citations2014Tim Salimans, Diederik P. Kingma +1 more
A new synthesis of variational inference and Monte Carlo methods where one or more steps of MCMC is incorporated into the authors' variational approximation, resulting in a rich class of inference algorithms bridging the gap between variational methods and MCMC.
arXiv (Cornell University)Learning Generative Models with Visual Attention
6 Citations2013Yichuan Tang, Nitish Srivastava +1 more
