What you can cram into a single $&!#* vector: Probing sentence embeddings for linguistic properties
Published 1 January 2018Open access
Alexis Conneau, Germán Kruszewski, Guillaume Lample, Loïc Barrault, Marco Baroni
Citations582
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
10 probing tasks designed to capture simple linguistic features of sentences are introduced and used to study embeddings generated by three different encoders trained in eight distinct ways, uncovering intriguing properties of bothencoders and training methods.
Abstract
Comunicació presentada a: 56th Annual Meeting of the Association for Computational Linguistics celebrat del 15 al 20 de juliol de 2018 a Melbourne, Australia.
Keywords
Computer Science
Glove: Global Vectors for Word Representation
33,769 Citations2014Jeffrey Pennington, Richard Socher +1 more
A new global logbilinear regression model that combines the advantages of the two major model families in the literature: global matrix factorization and local context window methods and produces a vector space with meaningful substructure.
Learning Phrase Representations using RNN Encoder–Decoder for Statistical Machine Translation
24,447 Citations2014Kyunghyun Cho, Bart van Merriënboer +5 more
Qualitatively, the proposed RNN Encoder‐Decoder model learns a semantically and syntactically meaningful representation of linguistic phrases.
arXiv (Cornell University)Efficient Estimation of Word Representations in Vector Space
18,050 Citations2013Tomáš Mikolov, Kai Chen +2 more
Two novel model architectures for computing continuous vector representations of words from very large data sets are proposed and it is shown that these vectors provide state-of-the-art performance on the authors' test set for measuring syntactic and semantic word similarities.
arXiv (Cornell University)Sequence to Sequence Learning with Neural Networks
13,362 Citations2014Ilya Sutskever, Oriol Vinyals +1 more
arXiv (Cornell University)Efficient Estimation of Word Representations in Vector Space
11,710 Citations2013Tomáš Mikolov, Kai Chen +2 more
A unified architecture for natural language processing
5,204 Citations2008Ronan Collobert, Jason Weston
This work describes a single convolutional neural network architecture that, given a sentence, outputs a host of language processing predictions: part-of-speech tags, chunks, named entity tags, semantic roles, semantically similar words and the likelihood that the sentence makes sense using a language model.
Moses
4,868 Citations2007Philipp Koehn, Richard Zens +12 more
An open-source toolkit for statistical machine translation whose novel contributions are support for linguistically motivated factors, confusion network decoding, and efficient data formats for translation models and language models.
A large annotated corpus for learning natural language inference
3,452 Citations2015Samuel R. Bowman, Gabor Angeli +2 more
The Stanford Natural Language Inference corpus is introduced, a new, freely available collection of labeled sentence pairs, written by humans doing a novel grounded task based on image captioning, which allows a neural network-based model to perform competitively on natural language inference benchmarks for the first time.
A sentimental education
3,343 Citations2004Bo Pang, Lillian Lee
A novel machine-learning method is proposed that applies text-categorization techniques to just the subjective portions of the document, which greatly facilitates incorporation of cross-sentence contextual constraints.
Europarl: A Parallel Corpus for Statistical Machine Translation
3,106 Citations2005Philipp Koehn
A corpus of parallel text in 11 languages from the proceedings of the European Parliament is collected and its acquisition and application as training data for statistical machine translation (SMT) is focused on.
Accurate unlexicalized parsing
3,055 Citations2003Dan Klein, Christopher D. Manning
It is demonstrated that an unlexicalized PCFG can parse much more accurately than previously shown, by making use of simple, linguistically motivated state splits, which break down false independence assumptions latent in a vanilla treebank grammar.
Linguistic Regularities in Continuous Space Word Representations
2,882 Citations2013Tomáš Mikolov, Wen-tau Yih +1 more
The vector-space word representations that are implicitly learned by the input-layer weights are found to be surprisingly good at capturing syntactic and semantic regularities in language, and that each relationship is characterized by a relation-specific vector offset.
arXiv (Cornell University)Supervised Learning of Universal Sentence Representations from Natural\n Language Inference Data
2,052 Citations2017Alexis Conneau, Douwe Kiela +3 more
Aligning Books and Movies: Towards Story-Like Visual Explanations by Watching Movies and Reading Books
2,035 Citations2015Yukun Zhu, Ryan Kiros +5 more
To align movies and books, a neural sentence embedding that is trained in an unsupervised way from a large corpus of books, as well as a video-text neural embedding for computing similarities between movie clips and sentences in the book are proposed.
arXiv (Cornell University)Convolutional Sequence to Sequence Learning
1,897 Citations2017Jonas Gehring, Michael Auli +3 more
This work introduces an architecture based entirely on convolutional neural networks, which outperform the accuracy of the deep LSTM setup of Wu et al. (2016) on both WMT'14 English-German and WMT-French translation at an order of magnitude faster speed, both on GPU and CPU.
RWTH Publications (RWTH Aachen)Moses: Open Source Toolkit for Statistical Machine Translation
1,451 Citations2007Philipp Koehn, Hieu Hoang +12 more
International Conference on Machine LearningConvolutional Sequence to Sequence Learning
1,340 Citations2017Jonas Gehring, Michael Auli +3 more
arXiv (Cornell University)Language Modeling with Gated Convolutional Networks
1,129 Citations2016Yann Dauphin, Angela Fan +2 more
International Conference on Learning RepresentationsA Simple but Tough-to-Beat Baseline for Sentence Embeddings
1,054 Citations2017Sanjeev Arora, Yingyu Liang +1 more
International Conference on Machine LearningLanguage modeling with gated convolutional networks
884 Citations2017Yann Dauphin, Angela Fan +2 more
Dynamic Pooling and Unfolding Recursive Autoencoders for Paraphrase Detection
810 Citations2011Richard Socher, Eric Huang +3 more
This work introduces a method for paraphrase detection based on recursive autoencoders (RAE) and unsupervised RAEs based on a novel unfolding objective and learns feature vectors for phrases in syntactic trees to measure word- and phrase-wise similarity between two sentences.
Transactions of the Association for Computational LinguisticsAssessing the Ability of LSTMs to Learn Syntax-Sensitive Dependencies
783 Citations2016Tal Linzen, Emmanuel Dupoux +1 more
It is concluded that LSTMs can capture a non-trivial amount of grammatical structure given targeted supervision, but stronger architectures may be required to further reduce errors; furthermore, the language modeling signal is insufficient for capturing syntax-sensitive dependencies, and should be supplemented with more direct supervision if such dependencies need to be captured.
International Journal of Computer VisionDeep Image Prior
715 Citations2020Dmitry Ulyanov, Andrea Vedaldi +1 more
It is shown that a randomly-initialized neural network can be used as a handcrafted prior with excellent results in standard inverse problems such as denoising, super-resolution, and inpainting.
A SICK cure for the evaluation of compositional distributional semantic models
661 Citations2014Marco Marelli, Stefano Menini +4 more
This work aims to help the research community working on compositional distributional semantic models (CDSMs) by providing SICK (Sentences Involving Compositional Knowldedge), a large size English benchmark tailored for them.
ArXiv.orgA Sentimental Education: Sentiment Analysis Using Subjectivity Summarization Based on Minimum Cuts
660 Citations2004Bo Pang, Lillian Lee
Visualizing and Understanding Neural Models in NLP
535 Citations2016Jiwei Li, Xinlei Chen +2 more
Four strategies for visualizing compositionality in neural models for NLP, inspired by similar work in computer vision, including LSTM-style gates that measure information flow and gradient back-propagation, are described.
arXiv (Cornell University)Grammar as a Foreign Language
403 Citations2014Oriol Vinyals, Łukasz Kaiser +4 more
Proceedings of the National Academy of SciencesNeurophysiological dynamics of phrase-structure building during sentence processing
348 Citations2017Matthew J. Nelson, Imen El Karoui +9 more
The results provide initial intracranial evidence for the neurophysiological reality of the merge operation postulated by linguists and suggest that the brain compresses syntactically well-formed sequences of words into a hierarchy of nested phrases.
arXiv (Cornell University)Fine-grained Analysis of Sentence Embeddings Using Auxiliary Prediction Tasks
295 Citations2016Yossi Adi, Einat Kermany +3 more
This work proposes a framework that facilitates better understanding of the encoded representations of sentence vectors and demonstrates the potential contribution of the approach by analyzing different sentence representation mechanisms.
What do Neural Machine Translation Models Learn about Morphology?
279 Citations2017Yonatan Belinkov, Nadir Durrani +3 more
This work analyzes the representations learned by neural MT models at various levels of granularity and empirically evaluates the quality of the representations for learning morphology through extrinsic part-of-speech and morphological tagging tasks.
arXiv (Cornell University)Fine-grained Analysis of Sentence Embeddings Using Auxiliary Prediction Tasks
241 Citations2024Yossi Adi
Transactions of the Association for Computational LinguisticsDeep Recurrent Models with Fast-Forward Connections for Neural Machine Translation
227 Citations2016Jie Zhou, Ying Cao +3 more
This work introduces a new type of linear connections, named fast-forward connections, based on deep Long Short-Term Memory (LSTM) networks, and an interleaved bi-directional architecture for stacking the LSTM layers, and achieves state-of-the-art performance and outperforms the best conventional model by 0.7 BLEU points.
Lecture notes in computer scienceRevisiting Visual Question Answering Baselines
225 Citations2016Allan Jabri, Armand Joulin +1 more
arXiv (Cornell University)Visualizing and Understanding Neural Models in NLP
166 Citations2015Jiwei Li, Xinlei Chen +2 more
Illinois-LH: A Denotational and Distributional Approach to Semantics
142 Citations2014Alice Lai, Julia Hockenmaier
This paper describes and analyzes the SemEval 2014 Task 1 system, which features are based on distributional and denotational similarities; word alignment; negation; and hypernym/hyponym, synonym, and antonym relations.
Computational LinguisticsRepresentation of Linguistic Form and Function in Recurrent Neural Networks
137 Citations2017Ákos Kádár, Grzegorz Chrupała +1 more
A method for estimating the amount of contribution of individual tokens in the input to the final prediction of the networks is proposed and shows that the Visual pathway pays selective attention to lexical categories and grammatical functions that carry semantic information, and learns to treat word types differently depending on their grammatical function and their position in the sequential structure of the sentence.
arXiv (Cornell University)Learning General Purpose Distributed Sentence Representations via Large\n Scale Multi-task Learning
127 Citations2018Sandeep Subramanian, Adam Trischler +2 more
International Joint Conference on Natural Language ProcessingEvaluating Layers of Representation in Neural Machine Translation on Part-of-Speech and Semantic Tagging Tasks
100 Citations2017Yonatan Belinkov, Lluı́s Màrquez +4 more
arXiv (Cornell University)Evaluating Layers of Representation in Neural Machine Translation on Part-of-Speech and Semantic Tagging Tasks
81 Citations2018Yonatan Belinkov, Lluı́s Màrquez +4 more
This paper investigates the quality of vector representations learned at different layers of NMT encoders and finds that higher layers are better at learning semantics while lower layers tend to be better for part-of-speech tagging.
Exploring how deep neural networks form phonemic categories
80 Citations2015Tasha Nagamine, Michael L. Seltzer +1 more
Phonetic features organize the activations in different layers of a DNN, a result that mirrors the recent findings of feature encoding in the human auditory system and may provide better understanding of the limitations of current models, leading to new strategies to improve their performance.
International Joint Conference on Natural Language ProcessingUnderstanding and Improving Morphological Learning in the Neural Machine Translation Decoder
57 Citations2017Fahim Dalvi, Nadir Durrani +3 more
This paper analyzes how much morphology an NMT decoder learns, and investigates whether injecting target morphology in the decoder helps it to produce better translations, and presents three methods for simultaneous translation, joint-data learning, and multi-task learning.
Jointly optimizing word representations for lexical and sentential tasks with the C-PHRASE model
50 Citations2015Nghia The Pham, Germán Kruszewski +2 more
C-PHRASE, a distributional semantic model that learns word representations by optimizing context prediction for phrases at all levels in a syntactic tree, outperforms the state-of-theart C-BOW model on a variety of lexical tasks.
Journal of Artificial Intelligence ResearchVisualisation and 'Diagnostic Classifiers' Reveal How Recurrent and Recursive Neural Networks Process Hierarchical Structure
9 Citations2018Dieuwke Hupkes, Sara Veldhoen +1 more
It is argued that diagnostic classification, unlike most visualisation techniques, does scale up from small networks in a toy domain, to larger and deeper recurrent networks dealing with real-life data, and may therefore contribute to a better understanding of the internal dynamics of current state-of-the-art models in natural language processing.
arXiv (Cornell University)Trimming and Improving Skip-thought Vectors
6 Citations2017Shuai Tang, Hailin Jin +3 more
It is found that a good word embedding initialization is also essential for learning better sentence representations, and the proposed model is a faster, lighter-weight and equally powerful alternative to the original skip-thought model.
