Neural Network Methods for Natural Language Processing
Synthesis lectures on human language technologiesPublished 1 January 2017
Yoav Goldberg
Citations771
SJR quartileQ3
SJR score0.12
SNIP0.00
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
This book focuses on the application of neural network models to natural language data, and introduces more specialized neural network architectures, including 1D convolutional neural networks, recurrent neural Networks, conditioned-generation models, and attention-based models.
Abstract
Neural networks are a family of powerful machine learning models. This book focuses on the application of neural network models to natural language data. The first half of the book (Parts I and II) co
Keywords
Computer Science
Deep Residual Learning for Image Recognition
222,082 Citations2016Kaiming He, Xiangyu Zhang +2 more
This work presents a residual learning framework to ease the training of networks that are substantially deeper than those used previously, and provides comprehensive empirical evidence showing that these residual networks are easier to optimize, and can gain accuracy from considerably increased depth.
Neural ComputationLong Short-Term Memory
98,079 Citations1997Sepp Hochreiter, Jürgen Schmidhuber
A novel, efficient, gradient based method called long short-term memory (LSTM) is introduced, which can learn to bridge minimal time lags in excess of 1000 discrete-time steps by enforcing constant error flow through constant error carousels within special units.
Proceedings of the IEEEGradient-based learning applied to document recognition
58,219 Citations1998Yann LeCun, Léon Bottou +2 more
This paper reviews various methods applied to handwritten character recognition and compares them on a standard handwritten digit recognition task, and Convolutional neural networks are shown to outperform all other techniques.
Dropout: a simple way to prevent neural networks from overfitting
34,279 Citations2014Nitish Srivastava, Geoffrey E. Hinton +3 more
It is shown that dropout improves the performance of neural networks on supervised learning tasks in vision, speech recognition, document classification and computational biology, obtaining state-of-the-art results on many benchmark data sets.
NatureLearning representations by back-propagating errors
30,885 Citations1986David E. Rumelhart, Geoffrey E. Hinton +1 more
Back-propagation repeatedly adjusts the weights of the connections in the network so as to minimize a measure of the difference between the actual output vector of the net and the desired output vector, which helps to represent important features of the task domain.
Neural NetworksMultilayer feedforward networks are universal approximators
21,248 Citations1989Kurt Hornik, Maxwell B. Stinchcombe +1 more
It is rigorously established that standard multilayer feedforward networks with as few as one hidden layer using arbitrary squashing functions are capable of approximating any Borel measurable function from one finite dimensional space to another to any desired degree of accuracy, provided sufficiently many hidden units are available.
Journal of the Royal Statistical Society Series B (Statistical Methodology)Regularization and Variable Selection Via the Elastic Net
20,982 Citations2005Hui Zou, Trevor Hastie
Delving Deep into Rectifiers: Surpassing Human-Level Performance on ImageNet Classification
18,777 Citations2015Kaiming He, Xiangyu Zhang +2 more
This work proposes a Parametric Rectified Linear Unit (PReLU) that generalizes the traditional rectified unit and derives a robust initialization method that particularly considers the rectifier nonlinearities.
Mathematics of Control Signals and SystemsApproximation by superpositions of a sigmoidal function
13,684 Citations1989George Cybenko
A training algorithm for optimal margin classifiers
11,594 Citations1992Bernhard E. Boser, Isabelle Guyon +1 more
A training algorithm that maximizes the margin between the training patterns and the decision boundary is presented, applicable to a wide variety of the classification functions, including Perceptrons, polynomials, and Radial Basis Functions.
Cognitive ScienceFinding Structure in Time
10,841 Citations1990Jeffrey L. Elman
A proposal along these lines first described by Jordan (1986) which involves the use of recurrent links in order to provide networks with a dynamic memory and suggests a method for representing lexical categories and the type/token distinction is developed.
IEEE Transactions on Signal ProcessingBidirectional recurrent neural networks
10,022 Citations1997Mike Schuster, Kuldip K. Paliwal
It is shown how the proposed bidirectional structure can be easily modified to allow efficient estimation of the conditional posterior probability of complete symbol sequences without making any explicit assumption about the shape of the distribution.
IEEE Transactions on Neural Networks and Learning SystemsLSTM: A Search Space Odyssey
6,837 Citations2016Klaus Greff, Rupesh K. Srivastava +3 more
This paper presents the first large-scale analysis of eight LSTM variants on three representative tasks: speech recognition, handwriting recognition, and polyphonic music modeling, and observes that the studied hyperparameters are virtually independent and derive guidelines for their efficient adjustment.
Show and tell: A neural image caption generator
6,363 Citations2015Oriol Vinyals, Alexander Toshev +2 more
This paper presents a generative model based on a deep recurrent architecture that combines recent advances in computer vision and machine translation and that can be used to generate natural sentences describing an image.
A unified architecture for natural language processing
5,204 Citations2008Ronan Collobert, Jason Weston
This work describes a single convolutional neural network architecture that, given a sentence, outputs a host of language processing predictions: part-of-speech tags, chunks, named entity tags, semantic roles, semantically similar words and the likelihood that the sentence makes sense using a language model.
Deep visual-semantic alignments for generating image descriptions
4,978 Citations2015Andrej Karpathy, Li Fei-Fei
A model that generates natural language descriptions of images and their regions using a novel combination of Convolutional Neural Networks over image regions, bidirectional Recurrent Neural Networks over sentences, and a structured objective that aligns the two modalities through a multimodal embedding is presented.
Curriculum learning
4,978 Citations2009Yoshua Bengio, Jérôme Louradour +2 more
It is hypothesized that curriculum learning has both an effect on the speed of convergence of the training process to a minimum and on the quality of the local minima obtained: curriculum learning can be seen as a particular form of continuation method (a general strategy for global optimization of non-convex functions).
Proceedings of the IEEEBackpropagation through time: what it does and how to do it
4,904 Citations1990Paul J. Werbos
This paper first reviews basic backpropagation, a simple method which is now being widely used in areas like pattern recognition and fault diagnosis, and describes further extensions of this method, to deal with systems other than neural networks, systems involving simultaneous equations or true recurrent networks, and other practical issues which arise with this method.
PsychometrikaThe Approximation of One Matrix by Another of Lower Rank
3,795 Citations1936Carl Eckart, Gale Young
Journal of the Royal Statistical Society Series B (Statistical Methodology)Regression Shrinkage and Selection via The Lasso: A Retrospective
3,669 Citations2011Robert Tibshirani
Studies in computational intelligenceSupervised Sequence Labelling with Recurrent Neural Networks
3,084 Citations2012Alex Graves
A new type of output layer that allows recurrent networks to be trained directly for sequence labelling tasks where the alignment between the inputs and the labels is unknown, and an extension of the long short-term memory network architecture to multidimensional data, such as images and video sequences.
Journal of Artificial Intelligence ResearchFrom Frequency to Meaning: Vector Space Models of Semantics
2,883 Citations2010Peter D. Turney, Patrick Pantel
The goal in this survey is to show the breadth of applications of VSMs for semantics, to provide a new perspective on VSMs, and to provide pointers into the literature for those who are less familiar with the field.
USSR Computational Mathematics and Mathematical PhysicsSome methods of speeding up the convergence of iteration methods
2,823 Citations1964B. T. Polyak
Computer Speech & LanguageAn empirical study of smoothing techniques for language modeling
2,077 Citations1999Stanley F. Chen, Joshua Goodman
Discriminative training methods for hidden Markov models
1,889 Citations2002Michael Collins
Experimental results on part-of-speech tagging and base noun phrase chunking are given, in both cases showing improvements over results for a maximum-entropy tagger.
Lecture notes in computer sciencePractical Recommendations for Gradient-Based Training of Deep Architectures
1,868 Citations2012Yoshua Bengio
Overall, this chapter describes elements of the practice used to successfully and efficiently train and debug large-scale and often deep multi-layer neural networks and closes with open questions about the training difficulties observed with deeper architectures.
Lecture notes in computer scienceStochastic Gradient Descent Tricks
1,860 Citations2012Léon Bottou
This chapter provides background material, explains why SGD is a good learning algorithm when the training set is large, and provides useful recommendations.
Lecture notes in computer scienceThe PASCAL Recognising Textual Entailment Challenge
1,626 Citations2006Ido Dagan, Oren Glickman +1 more
Extensions of recurrent neural network language model
1,596 Citations2011Tomáš Mikolov, Stefan Kombrink +3 more
Several modifications of the original recurrent neural network language model are presented, showing approaches that lead to more than 15 times speedup for both training and testing phases and possibilities how to reduce the amount of parameters in the model.
Improved backing-off for M-gram language modeling
1,490 Citations2002Reinhard Kneser, Hermann Ney
This paper proposes to use distributions which are especially optimized for the task of back-off, which are quite different from the probability distributions that are usually used for backing-off.
Design challenges and misconceptions in named entity recognition
1,476 Citations2009Lev Ratinov, Dan Roth
Some of the fundamental design challenges and misconceptions that underlie the development of an efficient and robust NER system are analyzed, and several solutions to these challenges are developed.
Transactions of the Association for Computational LinguisticsImproving Distributional Similarity with Lessons Learned from Word Embeddings
1,342 Citations2015Omer Levy, Yoav Goldberg +1 more
It is revealed that much of the performance gains of word embeddings are due to certain system design choices and hyperparameter optimizations, rather than the embedding algorithms themselves, and these modifications can be transferred to traditional distributional models, yielding similar gains.
Lecture notes in computer scienceMining the Web for Synonyms: PMI-IR versus LSA on TOEFL
1,298 Citations2001Peter D. Turney
Improving deep neural networks for LVCSR using rectified linear units and dropout
1,270 Citations2013George E. Dahl, Tara N. Sainath +1 more
Modelling deep neural networks with rectified linear unit (ReLU) non-linearities with minimal human hyper-parameter tuning on a 50-hour English Broadcast News task shows an 4.2% relative improvement over a DNN trained with sigmoid units, and a 14.4% relative improved over a strong GMM/HMM system.
Word association norms, mutual information, and lexicography
1,066 Citations1989Kenneth Church, Patrick Hanks
The proposed measure, the association ratio, estimates word association norms directly from computer readable corpora, making it possible to estimate norms for tens of thousands of words.
Feature hashing for large scale multitask learning
937 Citations2009Kilian Q. Weinberger, Anirban Dasgupta +3 more
This paper provides exponential tail bounds for feature hashing and shows that the interaction between random subspaces is negligible with high probability, and demonstrates the feasibility of this approach with experimental results for a new use case --- multitask learning.
Coarse-to-fine <i>n</i>-best parsing and MaxEnt discriminative reranking
888 Citations2005Eugene Charniak, Mark Johnson
This paper describes a simple yet novel method for constructing sets of 50- best parses based on a coarse-to-fine generative parser that generates 50-best lists that are of substantially higher quality than previously obtainable.
Artificial IntelligenceRecursive distributed representations
866 Citations1990Jordan Pollack
This paper presents a connectionist architecture which automatically develops compact distributed representations for variable-sized recursive data structures, as well as efficient accessing mechanisms for them.
Behavior Research MethodsExtracting semantic representations from word co-occurrence statistics: A computational study
819 Citations2007John A. Bullinaria, Joseph P. Levy
This article presents a systematic exploration of the principal computational possibilities for formulating and validating representations of word meanings from word co-occurrence statistics and finds that, once the best procedures are identified, a very simple approach is surprisingly successful and robust over a range of psychologically relevant evaluation measures.
Online large-margin training of dependency parsers
816 Citations2005Ryan McDonald, Koby Crammer +1 more
An effective training algorithm for linearly-scored dependency parsers that implements online large-margin multi-class training on top of efficient parsing techniques for dependency trees is presented.
Transactions of the Association for Computational LinguisticsAssessing the Ability of LSTMs to Learn Syntax-Sensitive Dependencies
783 Citations2016Tal Linzen, Emmanuel Dupoux +1 more
It is concluded that LSTMs can capture a non-trivial amount of grammatical structure given targeted supervision, but stronger architectures may be required to further reduce errors; furthermore, the language modeling signal is insufficient for capturing syntax-sensitive dependencies, and should be supplemented with more direct supervision if such dependencies need to be captured.
Computational LinguisticsDistributional Memory: A General Framework for Corpus-Based Semantics
662 Citations2010Marco Baroni, Alessandro Lenci
The Distributional Memory approach is shown to be tenable despite the constraints imposed by its multi-purpose nature, and performs competitively against task-specific algorithms recently reported in the literature for the same tasks, and against several state-of-the-art methods.
Computer Speech & LanguageA maximum entropy approach to adaptive statistical language modelling
548 Citations1996Roni Rosenfeld
An adaptive statistical language model is described, which successfully integrates long distance linguistic information with other knowledge sources, and shows the feasibility of incorporating many diverse knowledge sources in a single, unified statistical framework.
Computer Speech & LanguageA bit of progress in language modeling
463 Citations2001Joshua Goodman
A combination of all techniques together to a Katz smoothed trigram model with no count cutoffs achieves perplexity reductions between 38 and 50% (1 bit of entropy), depending on training data size, as well as a word error rate reduction of 8.9%.
Publications (Konstfack University of Arts, Crafts, and Design)The Distributional Hypothesis
459 Citations2008Magnus Sahlgren
There is a correlation between distributional similarity and meaning similarity, which allows us to utilize the former in order to estimate the latter, and one can pose two very basic questions concerning the distributional hypothesis: what kind of distributional properties the authors should look for, and what — if any — the differences are between different kinds of Distributional properties.
Computational LinguisticsAlgorithms for Deterministic Incremental Dependency Parsing
449 Citations2008Joakim Nivre
This article presents a general framework for describing and analyzing algorithms for deterministic incremental dependency parsing, formalized as transition systems and shows that all four algorithms give competitive accuracy, although the non-projective list-based algorithm generally outperforms the projective algorithms for languages with a non-negligible proportion of non- projective constructions.
Machine LearningSearch-based structured prediction
436 Citations2009Hal Daumé, John Langford +1 more
Searn is an algorithm for integrating search and learning to solve complex structured prediction problems such as those that occur in natural language, speech, computational biology, and vision and comes with a strong, natural theoretical guarantee: good performance on the derived classification problems implies goodperformance on the structured prediction problem.
Computational LinguisticsDiscriminative Reranking for Natural Language Parsing
407 Citations2005Michael J. Collins, Terry Koo
The boosting approach to ranking problems described in Freund et al. (1998) is applied to parsing the Wall Street Journal treebank, and it is argued that the method is an appealing alternative-in terms of both simplicity and efficiency-to work on feature selection methods within log-linear (maximum-entropy) models.
ACM Transactions on Speech and Language ProcessingUnsupervised models for morpheme segmentation and morphology learning
337 Citations2007Mathias Creutz, Krista Lagus
Morfessor can handle highly inflecting and compounding languages where words can consist of lengthy sequences of morphemes and is shown to perform very well compared to a widely known benchmark algorithm on Finnish data.
Factored language models and generalized parallel backoff
283 Citations2003Jeff Bilmes, Katrin Kirchhoff
Initial perplexity results on both CallHome Arabic and on Penn Treebank Wall Street Journal articles are provided, and FLMs with GPB can produce bigrams with significantly lower perplexity, sometimes lower than highly-optimized baseline trigrams.
SIAM ReviewIntroduction to Automatic Differentiation and MATLAB Object-Oriented Programming
262 Citations2010Richard D. Neidinger
A survey of more advanced topics in automatic differentiation includes an introduction to the reverse mode (the authors' implementation is forward mode) and considerations in arbitrary-order multivariable series computation.
Trends in Cognitive SciencesStructures, Not Strings: Linguistics as Part of the Cognitive Sciences
225 Citations2015Martin Everaert, Marinus A. C. Huybregts +3 more
Here it is shown how recent developments in generative grammar, taking language as a computational cognitive mechanism seriously, allow us to address issues left unexplained in the increasingly popular surface-oriented approaches to language.
A high-performance semi-supervised learning method for text chunking
223 Citations2005Rie Kubota Ando, Tong Zhang
A novel semi-supervised method that employs a learning paradigm which is to find "what good classifiers are like" by learning from thousands of automatically generated auxiliary classification problems on unlabeled data, which produces performance higher than the previous best results.
Fast methods for kernel-based text analysis
207 Citations2003Taku Kudo, Yūji Matsumoto
A Basket Mining algorithm is extended to convert a kernel-based classifier into a simple and fast linear classifier, showing results that show that these new classifiers are about 30 to 300 times faster than the standard kernel- based classifiers.
Smart Reply
183 Citations2016Anjuli Kannan, Karol Kurach +9 more
This paper proposes and investigates a novel end-to-end method for automatically generating short email responses, called Smart Reply, which generates semantically diverse suggestions that can be used as complete email responses with just one tap on mobile.
Efficient parsing for bilexical context-free grammars and head automaton grammars
160 Citations1999Jason Eisner, Giorgio Satta
This work presents O(n4) parsing algorithms for two bilexical formalisms, improving the prior upper bounds of O( n5) by one step and an improved grammar constant by another.
Disambiguation of super parts of speech (or supertags)
133 Citations1994Aravind K. Joshi, B. Srinivas
This work presents techniques for disambiguating supertags using local information such as lexical preference and local lexical dependencies, and the performance results for various models of supertag disambIGuation such as unigram, trigram and dependency-based models.
Continuous space language models for statistical machine translation
127 Citations2006Holger Schwenk, Daniel Dchelotte +1 more
This work proposes to use a new statistical language model that is based on a continuous representation of the words in the vocabulary, which achieves consistent improvements in the BLEU score on the development and test data.
Similarity-based estimation of word cooccurrence probabilities
119 Citations1994Ido Dagan, Fernando Pereira +1 more
A probabilistic word association model based on distributional word similarity is described, and it is applied to improving probability estimates for unseen word bigrams in a variant of Katz's back-off model.
Transactions of the Association for Computational LinguisticsTraining Deterministic Parsers with Non-Deterministic Oracles
108 Citations2013Yoav Goldberg, Joakim Nivre
Experimental evaluation on a wide range of data sets clearly shows that using dynamic oracles to train greedy parsers gives substantial improvements in accuracy, unlike other techniques like beam search.
Lecture notes in computer scienceRecursive hetero-associative memories for translation
84 Citations1997Mikel L. Forcada, Ramón P. Ñeco
A modification of Pollack's RAAM is presented, called a Recursive Hetero-Associative Memory (RHAM), and it is shown that it is capable of learning simple translation tasks, by building a state-Space representation of each input string and unfolding it to obtain the corresponding output string.
SemEval-2007 task 06
76 Citations2007Ken Litkowski, Orin Hargraves
The SemEval-2007 task to disambiguate prepositions was designed as a lexical sample task, and the data generated in the task provides ample opportunitites for further investigations of preposition behavior.
Lecture notes in computer scienceFacial Ethnicity Classification with Deep Convolutional Neural Networks
40 Citations2016Wei Wang, Feixiang He +1 more
This paper tackles the problem of ethnicity classification via using Deep Convolution Neural Networks to extract features and classify them simultaneously, and demonstrates the effectiveness of the proposed method.
Transactions of the Association for Computational LinguisticsImproved CCG Parsing with Semi-supervised Supertagging
39 Citations2014Mike Lewis, Mark Steedman
This work shows how a state-of-the-art CCG parser can be enhanced, by predicting lexical categories using unsupervised vector-space embeddings of words, which leads to substantial improvements in dependency parsing results over the standard supervised CCG parsers.
IEEE International Conference on Neural NetworksRule refinement with recurrent neural networks
32 Citations2002C. Lee Giles, Christian W. Omlin
The results from training a recurrent neural network to recognize a known nontrivial randomly generated regular grammar show that not only do the networks preserve correct prior knowledge, but they are able to correct through training inserted prior knowledge which was wrong.
Learning Distributed Representations of Conceptual Knowledge and their Application to Script-based Story Processing
22 Citations1992Geunbae Lee, Margot Flowers +1 more
A symbolic/connectionist hybrid script-based story processing system dynasty is constructed which incorporates DSR learning and 6 script related processing modules and is able to learn similarity-based distributed representations of concepts and events in everyday scriptal experiences.
The Journal of the Acoustical Society of AmericaA study of English word category prediction based on neural networks
16 Citations1988Masami Nakamura, Kiyohiro Shikano
Two neural network models that can learn hidden linguistic structure are proposed that can easily be expanded from Bigram to N‐gram networks and are effective for calculating the efficiency of this system.
Transactions of the Association for Computational LinguisticsSparse Non-negative Matrix Language Modeling
5 Citations2016Joris Pelemans, Noam Shazeer +1 more
Results show that SNM language models trained with n-gram features are a close match for the well-established Kneser-Ney models, and the addition of skip-gram features yields a model that is in the same league as the state-of-the-art recurrent neural network language models.
