Musical Audio Synthesis Using Autoencoding Neural Nets
Goldsmiths (University of London)Published 14 September 2014Open access
Andy M. Sarroff, Michael A. Casey
Citations31
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
An interactive musi- cal audio synthesis system that uses feedforward artificial neural networks for musical audio synthesis, rather than discriminative or regression tasks, and allows one to interact directly with the parameters of the model and generate musical audio in real time.
Abstract
(Abstract to follow)
Keywords
Computer Science
IEEE Transactions on Pattern Analysis and Machine IntelligenceRepresentation Learning: A Review and New Perspectives
13,002 Citations2013Yoshua Bengio, Aaron Courville +1 more
Recent work in the area of unsupervised feature learning and deep learning is reviewed, covering advances in probabilistic models, autoencoders, manifold learning, and deep networks.
Stacked Denoising Autoencoders: Learning Useful Representations in a Deep Network with a Local Denoising Criterion
5,014 Citations2010Pascal Vincent, Hugo Larochelle +3 more
Biological CyberneticsAuto-association by multilayer perceptrons and singular value decomposition
1,357 Citations1988H. Bourlard, Y. Kamp
It is shown that, for auto-association, the nonlinearities of the hidden units are useless and that the optimal parameter values can be derived directly by purely linear techniques relying on singular value decomposition and low rank matrix approximation, similar in spirit to the well-known Karhunen-Loève transform.
Neural Information Processing SystemsUnsupervised feature learning for audio classification using convolutional deep belief networks
941 Citations2009Honglak Lee, Peter T. Pham +2 more
On rectified linear units for speech processing
418 Citations2013Matthew D. Zeiler, M. Ranzato +9 more
This work shows that it can improve generalization and make training of deep networks faster and simpler by substituting the logistic units with rectified linear units.
Zenodo (CERN European Organization for Nuclear Research)Learning Features From Music Audio With Deep Belief Networks.
259 Citations2010Philippe Hamel, Douglas Eck
This work presents a system that can automatically extract relevant features from audio for a given task by using a Deep Belief Network on Discrete Fourier Transforms of the audio to solve the task of genre recognition.
arXiv (Cornell University)Pylearn2: a machine learning research library
193 Citations2013Ian Goodfellow, David Warde-Farley +7 more
A brief history of the library, an overview of its basic philosophy, a summary of the Library's architecture, and a description of how the Pylearn2 community functions socially are given.
Rethinking Automatic Chord Recognition with Convolutional Neural Networks
118 Citations2012Eric J. Humphrey, Juan Pablo Bello
This work adopts a different perspective of the problem, where several seconds of pitch spectra are classified directly by a convolutional neural network, and achieves state of the art performance through this initial effort to chord recognition.
Ghent University Academic Bibliography (Ghent University)Audio-Based Music Classification With A Pretrained Convolutional Network.
88 Citations2011Sander Dieleman, Phil ́ emon Brakel +1 more
A convolutional network is built that is then trained to perform artist recognition, genre recognition and key detection, and it is found that the Convolutional approach improves accuracy for the genre Recognition and artist recognition tasks.
Neural NetworksOnline learning and generalization of parts-based image representations by non-negative sparse autoencoders
79 Citations2012Andre Lemme, Felix Reinhart +1 more
It is shown that non-negativity constrains the space of solutions such that overfitting is prevented and very similar encodings are found irrespective of the network initialization and size.
Learning a robust Tonnetz-space transform for automatic chord recognition
56 Citations2012Eric J. Humphrey, Taemin Cho +1 more
A novel, data-driven approach to learning a robust function that projects audio data into Tonnetz-space, a geometric representation of equal-tempered pitch intervals grounded in music theory that out-performs the classification accuracy of previous chroma representations.
Feature Learning In Dynamic Environments: Modeling The Acoustic Structure Of Musical Emotion.
29 Citations2012Erik M. Schmidt, Jeffrey J. Scott +1 more
This work seeks to employ regression-based deep belief networks to learn features directly from magnitude spectra as a basis for feature learning, taking into account the dynamic nature of music.
Learning Rhythm And Melody Features With Deep Belief Networks.
17 Citations2013Erik M. Schmidt, Youngmoo E. Kim
Non-Linear Semantic Embedding for Organizing Large Instrument Sample Libraries
12 Citations2011Eric J. Humphrey, Aron P. Glennon +1 more
Non-linear Semantic Embedding (NLSE) is presented, capable of projecting high-dimensional time-frequency representations of musical instrument samples into a low-dimensional, semantically-organized metric space.
Zenodo (CERN European Organization for Nuclear Research)Cross-Modal Sound Mapping Using Deep Learning
10 Citations2013Ohad Fried, Rebecca Fiebrink
This work presents a method for automatic feature extraction and cross-modal mapping using deep learning, which uses stacked autoencoders to learn a layered feature representation of the data.
