Neural Network Methods for Natural Language Processing
Synthesis lectures on human language technologiesPublished 17 April 2017
Yoav Goldberg
Citations656
SJR quartileQ3
SJR score0.12
SNIP0.00
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
This book focuses on the application of neural network models to natural language data, and introduces more specialized neural network architectures, including 1D convolutional neural networks, recurrent neural Networks, conditioned-generation models, and attention-based models.
Abstract
Neural networks are a family of powerful machine learning models. This book focuses on the application of neural network models to natural language data. The first half of the book (Parts I and II) co
Keywords
Computer Science
Deep Residual Learning for Image Recognition
222,082 Citations2016Kaiming He, Xiangyu Zhang +2 more
This work presents a residual learning framework to ease the training of networks that are substantially deeper than those used previously, and provides comprehensive empirical evidence showing that these residual networks are easier to optimize, and can gain accuracy from considerably increased depth.
Neural ComputationLong Short-Term Memory
98,079 Citations1997Sepp Hochreiter, Jürgen Schmidhuber
A novel, efficient, gradient based method called long short-term memory (LSTM) is introduced, which can learn to bridge minimal time lags in excess of 1000 discrete-time steps by enforcing constant error flow through constant error carousels within special units.
Proceedings of the IEEEGradient-based learning applied to document recognition
58,219 Citations1998Yann LeCun, Léon Bottou +2 more
This paper reviews various methods applied to handwritten character recognition and compares them on a standard handwritten digit recognition task, and Convolutional neural networks are shown to outperform all other techniques.
NatureLearning representations by back-propagating errors
30,885 Citations1986David E. Rumelhart, Geoffrey E. Hinton +1 more
Back-propagation repeatedly adjusts the weights of the connections in the network so as to minimize a measure of the difference between the actual output vector of the net and the desired output vector, which helps to represent important features of the task domain.
NeurocomputingAdvances in neural information processing systems 7
22,300 Citations1997Krzysztof J. Cios, Mark E. Shields
Neural NetworksMultilayer feedforward networks are universal approximators
21,248 Citations1989Kurt Hornik, Maxwell B. Stinchcombe +1 more
It is rigorously established that standard multilayer feedforward networks with as few as one hidden layer using arbitrary squashing functions are capable of approximating any Borel measurable function from one finite dimensional space to another to any desired degree of accuracy, provided sufficiently many hidden units are available.
Journal of the Royal Statistical Society Series B (Statistical Methodology)Regularization and Variable Selection Via the Elastic Net
20,982 Citations2005Hui Zou, Trevor Hastie
Delving Deep into Rectifiers: Surpassing Human-Level Performance on ImageNet Classification
18,777 Citations2015Kaiming He, Xiangyu Zhang +2 more
This work proposes a Parametric Rectified Linear Unit (PReLU) that generalizes the traditional rectified unit and derives a robust initialization method that particularly considers the rectifier nonlinearities.
Mathematics of Control Signals and SystemsApproximation by superpositions of a sigmoidal function
13,684 Citations1989George Cybenko
A training algorithm for optimal margin classifiers
11,594 Citations1992Bernhard E. Boser, Isabelle Guyon +1 more
A training algorithm that maximizes the margin between the training patterns and the decision boundary is presented, applicable to a wide variety of the classification functions, including Perceptrons, polynomials, and Radial Basis Functions.
Cambridge University Press eBooksIntroduction to Information Retrieval
10,873 Citations2008Christopher D. Manning, Prabhakar Raghavan +1 more
This textbook teaches classical and web information retrieval, including web search and the related areas of text classification and text clustering from basic concepts, making it perfect for introductory courses in information retrieval for advanced undergraduates and graduate students in computer science.
Cognitive ScienceFinding Structure in Time
10,841 Citations1990Jeffrey L. Elman
A proposal along these lines first described by Jordan (1986) which involves the use of recurrent links in order to provide networks with a dynamic memory and suggests a method for representing lexical categories and the type/token distinction is developed.
IEEE Transactions on Signal ProcessingBidirectional recurrent neural networks
10,022 Citations1997Mike Schuster, Kuldip K. Paliwal
It is shown how the proposed bidirectional structure can be easily modified to allow efficient estimation of the conditional posterior probability of complete symbol sequences without making any explicit assumption about the shape of the distribution.
Foundations of statistical natural language processing
9,996 Citations1999Christopher D. Manning, Hinrich Schütze
IEEE Transactions on Neural Networks and Learning SystemsLSTM: A Search Space Odyssey
6,837 Citations2016Klaus Greff, Rupesh K. Srivastava +3 more
This paper presents the first large-scale analysis of eight LSTM variants on three representative tasks: speech recognition, handwriting recognition, and polyphonic music modeling, and observes that the studied hyperparameters are virtually independent and derive guidelines for their efficient adjustment.
Show and tell: A neural image caption generator
6,363 Citations2015Oriol Vinyals, Alexander Toshev +2 more
This paper presents a generative model based on a deep recurrent architecture that combines recent advances in computer vision and machine translation and that can be used to generate natural sentences describing an image.
Foundations and Trends® in Information RetrievalOpinion Mining and Sentiment Analysis
6,204 Citations2008Bo Pang, Lillian Lee
This survey covers techniques and approaches that promise to directly enable opinion-oriented information-seeking systems and focuses on methods that seek to address the new challenges raised by sentiment-aware applications, as compared to those that are already present in more traditional fact-based analysis.
A unified architecture for natural language processing
5,204 Citations2008Ronan Collobert, Jason Weston
This work describes a single convolutional neural network architecture that, given a sentence, outputs a host of language processing predictions: part-of-speech tags, chunks, named entity tags, semantic roles, semantically similar words and the likelihood that the sentence makes sense using a language model.
Deep visual-semantic alignments for generating image descriptions
4,978 Citations2015Andrej Karpathy, Li Fei-Fei
A model that generates natural language descriptions of images and their regions using a novel combination of Convolutional Neural Networks over image regions, bidirectional Recurrent Neural Networks over sentences, and a structured objective that aligns the two modalities through a multimodal embedding is presented.
Curriculum learning
4,978 Citations2009Yoshua Bengio, Jérôme Louradour +2 more
It is hypothesized that curriculum learning has both an effect on the speed of convergence of the training process to a minimum and on the quality of the local minima obtained: curriculum learning can be seen as a particular form of continuation method (a general strategy for global optimization of non-convex functions).
Proceedings of the IEEEBackpropagation through time: what it does and how to do it
4,904 Citations1990Paul J. Werbos
This paper first reviews basic backpropagation, a simple method which is now being widely used in areas like pattern recognition and fault diagnosis, and describes further extensions of this method, to deal with systems other than neural networks, systems involving simultaneous equations or true recurrent networks, and other practical issues which arise with this method.
HAL (Le Centre pour la Communication Scientifique Directe)Convolutional networks for images, speech, and time series
4,359 Citations1998Yann LeCun, Yoshua Bengio
PsychometrikaThe Approximation of One Matrix by Another of Lower Rank
3,795 Citations1936Carl Eckart, Gale Young
Journal of the Royal Statistical Society Series B (Statistical Methodology)Regression Shrinkage and Selection via The Lasso: A Retrospective
3,669 Citations2011Robert Tibshirani
Speech and language processing
3,589 Citations2010Dan Jurafsky, James Martin +1 more
It is now clear that HAL’s creator, Arthur C. Clarke, was a little optimistic in predicting when an artificial agent such as HAL would be avail-able.
Cambridge University Press eBooksUnderstanding Machine Learning
2,928 Citations2014Shai Shalev‐Shwartz, Shai Ben-David
Journal of Artificial Intelligence ResearchFrom Frequency to Meaning: Vector Space Models of Semantics
2,883 Citations2010Peter D. Turney, Patrick Pantel
The goal in this survey is to show the breadth of applications of VSMs for semantics, to provide a new perspective on VSMs, and to provide pointers into the literature for those who are less familiar with the field.
USSR Computational Mathematics and Mathematical PhysicsSome methods of speeding up the convergence of iteration methods
2,823 Citations1964B. T. Polyak
Computer Speech & LanguageAn empirical study of smoothing techniques for language modeling
2,077 Citations1999Stanley F. Chen, Joshua Goodman
Discriminative training methods for hidden Markov models
1,889 Citations2002Michael Collins
Experimental results on part-of-speech tagging and base noun phrase chunking are given, in both cases showing improvements over results for a maximum-entropy tagger.
Lecture notes in computer sciencePractical Recommendations for Gradient-Based Training of Deep Architectures
1,868 Citations2012Yoshua Bengio
Overall, this chapter describes elements of the practice used to successfully and efficiently train and debug large-scale and often deep multi-layer neural networks and closes with open questions about the training difficulties observed with deeper architectures.
Lecture notes in computer scienceStochastic Gradient Descent Tricks
1,860 Citations2012Léon Bottou
This chapter provides background material, explains why SGD is a good learning algorithm when the training set is large, and provides useful recommendations.
Lecture notes in computer scienceThe PASCAL Recognising Textual Entailment Challenge
1,626 Citations2006Ido Dagan, Oren Glickman +1 more
Extensions of recurrent neural network language model
1,596 Citations2011Tomáš Mikolov, Stefan Kombrink +3 more
Several modifications of the original recurrent neural network language model are presented, showing approaches that lead to more than 15 times speedup for both training and testing phases and possibilities how to reduce the amount of parameters in the model.
Improved backing-off for M-gram language modeling
1,490 Citations2002Reinhard Kneser, Hermann Ney
This paper proposes to use distributions which are especially optimized for the task of back-off, which are quite different from the probability distributions that are usually used for backing-off.
Design challenges and misconceptions in named entity recognition
1,476 Citations2009Lev Ratinov, Dan Roth
Some of the fundamental design challenges and misconceptions that underlie the development of an efficient and robust NER system are analyzed, and several solutions to these challenges are developed.
Transactions of the Association for Computational LinguisticsImproving Distributional Similarity with Lessons Learned from Word Embeddings
1,342 Citations2015Omer Levy, Yoav Goldberg +1 more
It is revealed that much of the performance gains of word embeddings are due to certain system design choices and hyperparameter optimizations, rather than the embedding algorithms themselves, and these modifications can be transferred to traditional distributional models, yielding similar gains.
Lecture notes in computer scienceMining the Web for Synonyms: PMI-IR versus LSA on TOEFL
1,298 Citations2001Peter D. Turney
Improving deep neural networks for LVCSR using rectified linear units and dropout
1,270 Citations2013George E. Dahl, Tara N. Sainath +1 more
Modelling deep neural networks with rectified linear unit (ReLU) non-linearities with minimal human hyper-parameter tuning on a 50-hour English Broadcast News task shows an 4.2% relative improvement over a DNN trained with sigmoid units, and a 14.4% relative improved over a strong GMM/HMM system.
The MIT Press eBooksFoundations of Machine Learning
1,083 Citations2012
This chapter contains sections titled: 2.1 A Direct Approach to Machine Learning, 2.2 General Methods of Analysis, and A Foundation for the Study of Boosting Algorithms.
Word association norms, mutual information, and lexicography
1,066 Citations1989Kenneth Church, Patrick Hanks
The proposed measure, the association ratio, estimates word association norms directly from computer readable corpora, making it possible to estimate norms for tens of thousands of words.
Feature hashing for large scale multitask learning
937 Citations2009Kilian Q. Weinberger, Anirban Dasgupta +3 more
This paper provides exponential tail bounds for feature hashing and shows that the interaction between random subspaces is negligible with high probability, and demonstrates the feasibility of this approach with experimental results for a new use case --- multitask learning.
Coarse-to-fine <i>n</i>-best parsing and MaxEnt discriminative reranking
888 Citations2005Eugene Charniak, Mark Johnson
This paper describes a simple yet novel method for constructing sets of 50- best parses based on a coarse-to-fine generative parser that generates 50-best lists that are of substantially higher quality than previously obtainable.
Lecture notes in computer scienceEfficient BackProp
867 Citations1998Yann LeCun, Léon Bottou +2 more
Many undesirable behaviors of backprop can be avoided with tricks that are rarely exposed in serious technical publications, and this paper gives some of those tricks, and offers explanations of why they work.
Artificial IntelligenceRecursive distributed representations
866 Citations1990Jordan Pollack
This paper presents a connectionist architecture which automatically develops compact distributed representations for variable-sized recursive data structures, as well as efficient accessing mechanisms for them.
Cambridge University Press eBooksStatistical Machine Translation
850 Citations2009Philipp Koehn
This introductory text to statistical machine translation (SMT) provides all of the theories and methods needed to build a statistical machine translator, such as Google Language Tools and Babelfish, and the companion website provides open-source corpora and tool-kits.
Behavior Research MethodsExtracting semantic representations from word co-occurrence statistics: A computational study
819 Citations2007John A. Bullinaria, Joseph P. Levy
This article presents a systematic exploration of the principal computational possibilities for formulating and validating representations of word meanings from word co-occurrence statistics and finds that, once the best procedures are identified, a very simple approach is surprisingly successful and robust over a range of psychologically relevant evaluation measures.
Online large-margin training of dependency parsers
816 Citations2005Ryan McDonald, Koby Crammer +1 more
An effective training algorithm for linearly-scored dependency parsers that implements online large-margin multi-class training on top of efficient parsing techniques for dependency trees is presented.
Transactions of the Association for Computational LinguisticsAssessing the Ability of LSTMs to Learn Syntax-Sensitive Dependencies
783 Citations2016Tal Linzen, Emmanuel Dupoux +1 more
It is concluded that LSTMs can capture a non-trivial amount of grammatical structure given targeted supervision, but stronger architectures may be required to further reduce errors; furthermore, the language modeling signal is insufficient for capturing syntax-sensitive dependencies, and should be supplemented with more direct supervision if such dependencies need to be captured.
Computational LinguisticsDistributional Memory: A General Framework for Corpus-Based Semantics
662 Citations2010Marco Baroni, Alessandro Lenci
The Distributional Memory approach is shown to be tenable despite the constraints imposed by its multi-purpose nature, and performs competitively against task-specific algorithms recently reported in the literature for the same tasks, and against several state-of-the-art methods.
Multitask Learning
562 Citations1998Rich Caruana
The MIT Press eBooksPredicting Structured Data
533 Citations2007
This volume presents the state of the art in machine learning algorithms and theory in this novel field and discusses applications as diverse as machine translation, document markup, computational biology, and information extraction, providing a timely overview of an exciting field.
Journal of the American Society for Information Science and TechnologyComputational methods in authorship attribution
498 Citations2008Moshe Koppel, Jonathan Schler +1 more
This work presents a meta-modelling framework that automates the very labor-intensive and therefore time-heavy and therefore expensive and expensive process of manually annotating manuscripts to establish authorship attribution.
Computer Speech & LanguageA bit of progress in language modeling
463 Citations2001Joshua Goodman
A combination of all techniques together to a Katz smoothed trigram model with no count cutoffs achieves perplexity reductions between 38 and 50% (1 bit of entropy), depending on training data size, as well as a word error rate reduction of 8.9%.
Computational LinguisticsAlgorithms for Deterministic Incremental Dependency Parsing
449 Citations2008Joakim Nivre
This article presents a general framework for describing and analyzing algorithms for deterministic incremental dependency parsing, formalized as transition systems and shows that all four algorithms give competitive accuracy, although the non-projective list-based algorithm generally outperforms the projective algorithms for languages with a non-negligible proportion of non- projective constructions.
Machine LearningSearch-based structured prediction
436 Citations2009Hal Daumé, John Langford +1 more
Searn is an algorithm for integrating search and learning to solve complex structured prediction problems such as those that occur in natural language, speech, computational biology, and vision and comes with a strong, natural theoretical guarantee: good performance on the derived classification problems implies goodperformance on the structured prediction problem.
Computational LinguisticsDiscriminative Reranking for Natural Language Parsing
407 Citations2005Michael J. Collins, Terry Koo
The boosting approach to ranking problems described in Freund et al. (1998) is applied to parsing the Wall Street Journal treebank, and it is argued that the method is an appealing alternative-in terms of both simplicity and efficiency-to work on feature selection methods within log-linear (maximum-entropy) models.
ACM Transactions on Speech and Language ProcessingUnsupervised models for morpheme segmentation and morphology learning
337 Citations2007Mathias Creutz, Krista Lagus
Morfessor can handle highly inflecting and compounding languages where words can consist of lengthy sequences of morphemes and is shown to perform very well compared to a widely known benchmark algorithm on Finnish data.
Factored language models and generalized parallel backoff
283 Citations2003Jeff Bilmes, Katrin Kirchhoff
Initial perplexity results on both CallHome Arabic and on Penn Treebank Wall Street Journal articles are provided, and FLMs with GPB can produce bigrams with significantly lower perplexity, sometimes lower than highly-optimized baseline trigrams.
Research Portal (King's College London)Transcending Boundaries: Improvisation and Disability in Dance
263 Citations2019Sarah Whatley
SIAM ReviewIntroduction to Automatic Differentiation and MATLAB Object-Oriented Programming
262 Citations2010Richard D. Neidinger
A survey of more advanced topics in automatic differentiation includes an introduction to the reverse mode (the authors' implementation is forward mode) and considerations in arbitrary-order multivariable series computation.
Trends in Cognitive SciencesStructures, Not Strings: Linguistics as Part of the Cognitive Sciences
225 Citations2015Martin Everaert, Marinus A. C. Huybregts +3 more
Here it is shown how recent developments in generative grammar, taking language as a computational cognitive mechanism seriously, allow us to address issues left unexplained in the increasingly popular surface-oriented approaches to language.
A high-performance semi-supervised learning method for text chunking
223 Citations2005Rie Kubota Ando, Tong Zhang
A novel semi-supervised method that employs a learning paradigm which is to find "what good classifiers are like" by learning from thousands of automatically generated auxiliary classification problems on unlabeled data, which produces performance higher than the previous best results.
Fast methods for kernel-based text analysis
207 Citations2003Taku Kudo, Yūji Matsumoto
A Basket Mining algorithm is extended to convert a kernel-based classifier into a simple and fast linear classifier, showing results that show that these new classifiers are about 30 to 300 times faster than the standard kernel- based classifiers.
Smart Reply
183 Citations2016Anjuli Kannan, Karol Kurach +9 more
This paper proposes and investigates a novel end-to-end method for automatically generating short email responses, called Smart Reply, which generates semantically diverse suggestions that can be used as complete email responses with just one tap on mobile.
Recognizing textual entailment models and applications
180 Citations2013Ido Dagan
Synthesis lectures on human language technologiesDependency Parsing
170 Citations2009Sandra Kübler, Ryan McDonald +1 more
Efficient parsing for bilexical context-free grammars and head automaton grammars
160 Citations1999Jason Eisner, Giorgio Satta
This work presents O(n4) parsing algorithms for two bilexical formalisms, improving the prior upper bounds of O( n5) by one step and an improved grammar constant by another.
Disambiguation of super parts of speech (or supertags)
133 Citations1994Aravind K. Joshi, B. Srinivas
This work presents techniques for disambiguating supertags using local information such as lexical preference and local lexical dependencies, and the performance results for various models of supertag disambIGuation such as unigram, trigram and dependency-based models.
Studies in fuzziness and soft computingNeural Probabilistic Language Models
130 Citations2006Yoshua Bengio, Holger Schwenk +3 more
Continuous space language models for statistical machine translation
127 Citations2006Holger Schwenk, Daniel Dchelotte +1 more
This work proposes to use a new statistical language model that is based on a continuous representation of the words in the vocabulary, which achieves consistent improvements in the BLEU score on the development and test data.
Similarity-based estimation of word cooccurrence probabilities
119 Citations1994Ido Dagan, Fernando Pereira +1 more
A probabilistic word association model based on distributional word similarity is described, and it is applied to improving probability estimates for unseen word bigrams in a variant of Katz's back-off model.
Transactions of the Association for Computational LinguisticsTraining Deterministic Parsers with Non-Deterministic Oracles
108 Citations2013Yoav Goldberg, Joakim Nivre
Experimental evaluation on a wide range of data sets clearly shows that using dynamic oracles to train greedy parsers gives substantial improvements in accuracy, unlike other techniques like beam search.
Lecture notes in computer scienceRecursive hetero-associative memories for translation
84 Citations1997Mikel L. Forcada, Ramón P. Ñeco
A modification of Pollack's RAAM is presented, called a Recursive Hetero-Associative Memory (RHAM), and it is shown that it is capable of learning simple translation tasks, by building a state-Space representation of each input string and unfolding it to obtain the corresponding output string.
SemEval-2007 task 06
76 Citations2007Ken Litkowski, Orin Hargraves
The SemEval-2007 task to disambiguate prepositions was designed as a lexical sample task, and the data generated in the task provides ample opportunitites for further investigations of preposition behavior.
arXiv (Cornell University)Long Short-Term Memory Over Tree Structures
68 Citations2015Xiaodan Zhu, Parinaz Sobhani +1 more
This paper proposes to extend chain-structured long short-term memory to tree structures, in which a memory cell can reflect the history memories of multiple child cells or multiple descendant cells in a recursive process, and calls the model S-LSTM, which provides a principled way of considering long-distance interaction over hierarchies.
Synthesis lectures on human language technologiesSemi-Supervised Learning and Domain Adaptation in Natural Language Processing
50 Citations2013Anders Søgaard
Lecture notes in computer scienceFacial Ethnicity Classification with Deep Convolutional Neural Networks
40 Citations2016Wei Wang, Feixiang He +1 more
This paper tackles the problem of ethnicity classification via using Deep Convolution Neural Networks to extract features and classify them simultaneously, and demonstrates the effectiveness of the proposed method.
Transactions of the Association for Computational LinguisticsImproved CCG Parsing with Semi-supervised Supertagging
39 Citations2014Mike Lewis, Mark Steedman
This work shows how a state-of-the-art CCG parser can be enhanced, by predicting lexical categories using unsupervised vector-space embeddings of words, which leads to substantial improvements in dependency parsing results over the standard supervised CCG parsers.
IEEE International Conference on Neural NetworksRule refinement with recurrent neural networks
32 Citations2002C. Lee Giles, Christian W. Omlin
The results from training a recurrent neural network to recognize a known nontrivial randomly generated regular grammar show that not only do the networks preserve correct prior knowledge, but they are able to correct through training inserted prior knowledge which was wrong.
Synthesis lectures on human language technologiesLinguistic Structure Prediction
30 Citations2011Noah A. Smith
Synthesis lectures on human language technologiesLinguistic Fundamentals for Natural Language Processing
26 Citations2013Emily M. Bender
Learning Distributed Representations of Conceptual Knowledge and their Application to Script-based Story Processing
22 Citations1992Geunbae Lee, Margot Flowers +1 more
A symbolic/connectionist hybrid script-based story processing system dynasty is constructed which incorporates DSR learning and 6 script related processing modules and is able to learn similarity-based distributed representations of concepts and events in everyday scriptal experiences.
The Journal of the Acoustical Society of AmericaA study of English word category prediction based on neural networks
16 Citations1988Masami Nakamura, Kiyohiro Shikano
Two neural network models that can learn hidden linguistic structure are proposed that can easily be expanded from Bigram to N‐gram networks and are effective for calculating the efficiency of this system.
Synthesis lectures on human language technologiesSyntax-based Statistical Machine Translation
13 Citations2016Philip Williams, Rico Sennrich +2 more
Transactions of the Association for Computational LinguisticsSparse Non-negative Matrix Language Modeling
5 Citations2016Joris Pelemans, Noam Shazeer +1 more
Results show that SNM language models trained with n-gram features are a close match for the well-established Kneser-Ney models, and the addition of skip-gram features yields a model that is in the same league as the state-of-the-art recurrent neural network language models.
…
