Hierarchical Attention Networks for Document Classification
Published 1 January 2016Open access
Zichao Yang, Diyi Yang, Chris Dyer, Xiaodong He, Alex Smola, Eduard Hovy
Citations4,829
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
Experiments conducted on six large scale text classification tasks demonstrate that the proposed architecture outperform previous methods by a substantial margin.
Abstract
Zichao Yang, Diyi Yang, Chris Dyer, Xiaodong He, Alex Smola, Eduard Hovy. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2016.
Keywords
Computer Science
Neural ComputationLong Short-Term Memory
98,079 Citations1997Sepp Hochreiter, Jürgen Schmidhuber
A novel, efficient, gradient based method called long short-term memory (LSTM) is introduced, which can learn to bridge minimal time lags in excess of 1000 discrete-time steps by enforcing constant error flow through constant error carousels within special units.
Proceedings of the IEEEGradient-based learning applied to document recognition
58,219 Citations1998Yann LeCun, Léon Bottou +2 more
This paper reviews various methods applied to handwritten character recognition and compares them on a standard handwritten digit recognition task, and Convolutional neural networks are shown to outperform all other techniques.
arXiv (Cornell University)Distributed Representations of Words and Phrases and their Compositionality
18,086 Citations2013Tomáš Mikolov, Ilya Sutskever +3 more
This paper presents a simple method for finding phrases in text, and shows that learning good vector representations for millions of phrases is possible and describes a simple alternative to the hierarchical softmax called negative sampling.
arXiv (Cornell University)Neural Machine Translation by Jointly Learning to Align and Translate
14,565 Citations2014Dzmitry Bahdanau
Convolutional Neural Networks for Sentence Classification
13,798 Citations2014Yoon Kim
The CNN models discussed herein improve upon the state of the art on 4 out of 7 tasks, which include sentiment analysis and question classification, and are proposed to allow for the use of both task-specific and static vectors.
Lecture notes in computer scienceText categorization with Support Vector Machines: Learning with many relevant features
7,925 Citations1998Thorsten Joachims
SVMs achieve substantial improvements over the currently best performing methods and behave robustly over a variety of di-erent learning tasks, eliminating the need for manual parameter tuning.
arXiv (Cornell University)Show, Attend and Tell: Neural Image Caption Generation with Visual Attention
7,525 Citations2015Kelvin Xu, Jimmy Ba +6 more
An attention based model that automatically learns to describe the content of images is introduced that can be trained in a deterministic manner using standard backpropagation techniques and stochastically by maximizing a variational lower bound.
The Stanford CoreNLP Natural Language Processing Toolkit
7,238 Citations2014Christopher D. Manning, Mihai Surdeanu +4 more
The design and use of the Stanford CoreNLP toolkit is described, an extensible pipeline that provides core natural language analysis, and it is suggested that this follows from a simple, approachable design, straightforward interfaces, the inclusion of robust and good quality analysis components, and not requiring use of a large amount of associated baggage.
Recursive Deep Models for Semantic Compositionality Over a Sentiment Treebank
6,774 Citations2013Richard Socher, Alex Perelygin +5 more
A Sentiment Treebank that includes fine grained sentiment labels for 215,154 phrases in the parse trees of 11,855 sentences and presents new challenges for sentiment compositionality, and introduces the Recursive Neural Tensor Network.
now publishers, Inc. eBooksOpinion Mining and Sentiment Analysis
6,745 Citations2008Bo Pang, Lillian Lee
A Convolutional Neural Network for Modelling Sentences
3,561 Citations2014Nal Kalchbrenner, Edward Grefenstette +1 more
A convolutional architecture dubbed the Dynamic Convolutional Neural Network (DCNN) is described that is adopted for the semantic modelling of sentences and induces a feature graph over the sentence that is capable of explicitly capturing short and long-range relations.
Learning Word Vectors for Sentiment Analysis
3,299 Citations2011Andrew L. Maas, Raymond E. Daly +4 more
This work presents a model that uses a mix of unsupervised and supervised techniques to learn word vectors capturing semantic term--document information as well as rich sentiment content, and finds it out-performs several previously introduced methods for sentiment classification.
arXiv (Cornell University)Character-level Convolutional Networks for Text Classification
3,266 Citations2015Xiang Zhang, Junbo Zhao +1 more
This article constructed several large-scale datasets to show that character-level convolutional networks could achieve state-of-the-art or competitive results in text classification.
Proceedings of the AAAI Conference on Artificial IntelligenceRecurrent Convolutional Neural Networks for Text Classification
2,306 Citations2015Siwei Lai, Liheng Xu +2 more
A recurrent convolutional neural network is introduced for text classification without human-designed features to capture contextual information as far as possible when learning word representations, which may introduce considerably less noise compared to traditional window-based neural networks.
Stacked Attention Networks for Image Question Answering
2,056 Citations2016Zichao Yang, Xiaodong He +3 more
A multiple-layer SAN is developed in which an image is queried multiple times to infer the answer progressively, and the progress that the SAN locates the relevant visual clues that lead to the answer of the question layer-by-layer.
Oxford University Research Archive (ORA) (University of Oxford)Teaching Machines to Read and Comprehend
1,936 Citations2015Karl Moritz Hermann, Tomáš Kočiský +5 more
Document Modeling with Gated Recurrent Neural Network for Sentiment Classification
1,535 Citations2015Duyu Tang, Bing Qin +1 more
A neural network model is introduced to learn vector-based document representation in a unified, bottom-up fashion and dramatically outperforms standard recurrent neural network in document modeling for sentiment classification.
arXiv (Cornell University)Teaching Machines to Read and Comprehend
1,526 Citations2015Karl Moritz Hermann, Tomáš Kočiský +5 more
A new methodology is defined that resolves this bottleneck and provides large scale supervised reading comprehension data that allows a class of attention based deep neural networks that learn to read real documents and answer complex questions with minimal prior knowledge of language structure to be developed.
Learning Sentiment-Specific Word Embedding for Twitter Sentiment Classification
1,167 Citations2014Duyu Tang, Furu Wei +4 more
Three neural networks are developed to effectively incorporate the supervision from sentiment polarity of text (e.g. sentences or tweets) in their loss functions and the performance of SSWE is improved by concatenating SSWE with existing feature set.
A Bayesian Approach to Filtering Junk E-Mail
1,153 Citations1998Mehran Sahami, Susan Dumais +2 more
This work examines methods for the automated construction of filters to eliminate such unwanted messages from a user’s mail stream, and shows the efficacy of such filters in a real world usage scenario, arguing that this technology is mature enough for deployment.
Baselines and Bigrams: Simple, Good Sentiment and Topic Classification
967 Citations2012Sida Wang, Christopher D. Manning
It is shown that the inclusion of word bigram features gives consistent gains on sentiment analysis tasks, and a simple but novel SVM variant using NB log-count ratios as feature values consistently performs well across tasks and datasets.
Journal of Artificial Intelligence ResearchSentiment Analysis of Short Informal Texts
890 Citations2014Svetlana Kiritchenko, Xiaodan Zhu +1 more
A state-of-the-art sentiment analysis system that detects (a) the sentiment of short informal textual messages such as tweets and SMS (message-level task) and (b) the Sentiment of a word or a phrase within a message (term- level task).
Effective Use of Word Order for Text Categorization with Convolutional Neural Networks
776 Citations2015Rie Johnson, Tong Zhang
A straightforward adaptation of CNN from image to text, a simple but new variation which employs bag-of-word conversion in the convolution layer is proposed and an extension to combine multiple convolution layers is explored for higher accuracy.
A Latent Semantic Model with Convolutional-Pooling Structure for Information Retrieval
682 Citations2014Yelong Shen, Xiaodong He +3 more
A new latent semantic model that incorporates a convolutional-pooling structure over word sequences to learn low-dimensional, semantic vector representations for search queries and Web documents is proposed.
arXiv (Cornell University)A C-LSTM Neural Network for Text Classification
656 Citations2015Chunting Zhou, Chonglin Sun +2 more
C-LSTM is a novel and unified model for sentence representation and text classification that outperforms both CNN and LSTM and can achieve excellent performance on these tasks.
arXiv (Cornell University)Ask Me Anything: Dynamic Memory Networks for Natural Language Processing
631 Citations2015Ankit Kumar, Ozan İrsoy +7 more
The dynamic memory network (DMN), a neural network architecture which processes input sequences and questions, forms episodic memories, and generates relevant answers, is introduced.
Finding Function in Form: Compositional Character Models for Open Vocabulary Word Representation
568 Citations2015Ling Wang, Chris Dyer +6 more
A model for constructing vector representations of words by composing characters using bidirectional LSTMs that requires only a single vector per character type and a fixed set of parameters for the compositional model, which yields state- of-the-art results in language modeling and part-of-speech tagging.
A Hierarchical Neural Autoencoder for Paragraphs and Documents
518 Citations2015Jiwei Li, Thang Luong +1 more
This paper introduces an LSTM model that hierarchically builds an embedding for a paragraph from embeddings for sentences and words, then decodes this embedding to reconstruct the original paragraph and evaluates the reconstructed paragraph using standard metrics to show that neural models are able to encode texts in a way that preserve syntactic, semantic, and discourse coherence.
Jointly modeling aspects, ratings and sentiments for movie recommendation (JMARS)
454 Citations2014Qiming Diao, Minghui Qiu +4 more
A probabilistic model based on collaborative filtering and topic modeling is proposed that allows it to capture the interest distribution of users and the content distribution for movies; it provides a link between interest and relevance on a per-aspect basis and it allows us to differentiate between positive and negative sentiments on aPer-Aspect basis.
arXiv (Cornell University)Grammar as a Foreign Language
403 Citations2014Oriol Vinyals, Łukasz Kaiser +4 more
Modeling Interestingness with Deep Neural Networks
175 Citations2014Jianfeng Gao, Patrick Pantel +3 more
The results on large-scale, real-world datasets show that the semantics of documents are important for modeling interest-ingness and that the DSSM leads to significant quality improvement on both tasks, outperforming not only the classic document models that do not use semantics but also state-of-the-art topic models.
Hierarchical Recurrent Neural Network for Document Modeling
165 Citations2015Rui Lin, Shujie Liu +4 more
A novel hierarchical recurrent neural network language model (HRNNLM) for document modeling that integrates it as the sentence history information into the word level RNN to predict the word sequence with cross-sentence contextual information.
arXiv (Cornell University)Finding Function in Form: Compositional Character Models for Open Vocabulary Word Representation
131 Citations2015Ling Wang, Tiago Luís +6 more
