Published 1 January 2019Open access
Alan Akbik, Tanja Bergmann, Duncan A. J. Blythe, Kashif Rasul, Stefan Schweter, Roland Vollgraf
Citations333
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
The core idea of the FLAIR framework is to present a simple, unified interface for conceptually very different types of word and document embeddings, which effectively hides all embedding-specific engineering complexity and allows researchers to “mix and match” variousembeddings with little effort.
Abstract
Alan Akbik, Tanja Bergmann, Duncan Blythe, Kashif Rasul, Stefan Schweter, Roland Vollgraf. Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics (Demonstrations). 2019.
Keywords
Computer Science
Glove: Global Vectors for Word Representation
33,769 Citations2014Jeffrey Pennington, Richard Socher +1 more
A new global logbilinear regression model that combines the advantages of the two major model families in the literature: global matrix factorization and local context window methods and produces a vector space with meaningful substructure.
32,525 Citations2019Jacob Devlin, Ming‐Wei Chang +2 more
A new language representation model, BERT, designed to pre-train deep bidirectional representations from unlabeled text by jointly conditioning on both left and right context in all layers, which can be fine-tuned with just one additional output layer to create state-of-the-art models for a wide range of tasks.
arXiv (Cornell University)Distributed Representations of Words and Phrases and their Compositionality
18,086 Citations2013Tomáš Mikolov, Ilya Sutskever +3 more
This paper presents a simple method for finding phrases in text, and shows that learning good vector representations for millions of phrases is possible and describes a simple alternative to the hierarchical softmax called negative sampling.
Transactions of the Association for Computational LinguisticsEnriching Word Vectors with Subword Information
9,784 Citations2017Piotr Bojanowski, Édouard Grave +2 more
A new approach based on the skipgram model, where each word is represented as a bag of character n-grams, with words being represented as the sum of these representations, which achieves state-of-the-art performance on word similarity and analogy tasks.
Neural Architectures for Named Entity Recognition
4,428 Citations2016Guillaume Lample, Miguel Ballesteros +3 more
Comunicacio presentada a la 2016 Conference of the North American Chapter of the Association for Computational Linguistics, celebrada a San Diego (CA, EUA) els dies 12 a 17 of juny 2016.
Learning Word Vectors for Sentiment Analysis
3,299 Citations2011Andrew L. Maas, Raymond E. Daly +4 more
This work presents a model that uses a mix of unsupervised and supervised techniques to learn word vectors capturing semantic term--document information as well as rich sentiment content, and finds it out-performs several previously introduced methods for sentiment classification.
Transformer-XL: Attentive Language Models beyond a Fixed-Length Context
3,149 Citations2019Zihang Dai, Zhilin Yang +4 more
This work proposes a novel neural architecture Transformer-XL that enables learning dependency beyond a fixed length without disrupting temporal coherence, which consists of a segment-level recurrence mechanism and a novel positional encoding scheme.
arXiv (Cornell University)Supervised Learning of Universal Sentence Representations from Natural\n Language Inference Data
2,052 Citations2017Alexis Conneau, Douwe Kiela +3 more
Digital Access to Scholarship at Harvard (DASH) (Harvard University)Making a Science of Model Search: Hyperparameter Optimization in Hundreds of Dimensions for Vision Architectures
1,644 Citations2013James Bergstra, Daniel Yamins +1 more
This work proposes a meta-modeling approach to support automated hyperparameter optimization, with the goal of providing practical tools that replace hand-tuning with a reproducible and unbiased optimization process.
Learning question classifiers
1,276 Citations2002Xin Li, Dan Roth
A hierarchical classifier is learned that is guided by a layered semantic hierarchy of answer types, and eventually classifies questions into fine-grained classes.
International Conference on Computational LinguisticsContextual String Embeddings for Sequence Labeling
1,004 Citations2018Alan Akbik, Duncan A. J. Blythe +1 more
This paper proposes to leverage the internal states of a trained character language model to produce a novel type of word embedding which they refer to as contextual string embeddings, which are fundamentally model words as sequences of characters and are contextualized by their surrounding text.
Dissecting Contextual Word Embeddings: Architecture and Representation
421 Citations2018Matthew E. Peters, Mark E Neumann +2 more
There is a tradeoff between speed and accuracy, but all architectures learn high quality contextual representations that outperform word embeddings for four challenging NLP tasks, suggesting that unsupervised biLMs, independent of architecture, are learning much more about the structure of language than previously appreciated.
Results of the WNUT2017 Shared Task on Novel and Emerging Entity Recognition
366 Citations2017Leon Derczynski, Eric Nichols +2 more
The goal of this task is to provide a definition of emerging and of rare entities, and based on that, also datasets for detecting these entities and to evaluate the ability of participating entries to detect and classify novel and emerging named entities in noisy text.
Artificial IntelligenceLearning multilingual named entity recognition from Wikipedia
366 Citations2012Joel Nothman, Nicky Ringland +3 more
The approach outperforms other approaches to automatic ne annotation; competes with gold-standard training when tested on an evaluation corpus from a different source; and performs 10% better than newswire-trained models on manually-annotated Wikipedia text.
Information Processing & ManagementOverview of the Sixth Text REtrieval Conference (TREC-6)
315 Citations2000Ellen M. Voorhees, Donna Harman
The Text REtrieval Conference is a workshop series designed to encourage research on text retrieval for realistic applications by providing large test collections, uniform scoring procedures and a forum for organizations interested in comparing results.
Pooled Contextualized Embeddings for Named Entity Recognition
307 Citations2019Alan Akbik, Tanja Bergmann +1 more
This work proposes a method in which it dynamically aggregate contextualized embeddings of each unique string that the authors encounter and uses a pooling operation to distill a ”global” word representation from all contextualized instances.
Publication Server of the Institute for German Language (Institute for German Language)Overview of the GermEval 2018 Shared Task on the Identification of Offensive Language
237 Citations2019Michael Wiegand, Melanie Siegel +1 more
This pilot edition of the GermEval Shared Task on the Identification of Offensive Language deals with the classification of German tweets from Twitter and describes the process of extracting the raw-data for the data collection and the annotation schema.
Conference on Computational Natural Language LearningCoNLL 2018 Shared Task : Multilingual Parsing from Raw Text to Universal Dependencies
184 Citations2018Daniel Zeman, Jan Hajič +6 more
PropBank: Semantics of New Predicate Types
48 Citations2014Claire Bonial, Julia Bonn +3 more
This research focuses on expanding PropBank, a corpus annotated with predicate argument structures, with new predicate types; namely, noun, adjective and complex predicates, such as Light Verb Constructions, in order for PropBank to reach the same level of coverage and continue to serve as the bedrock for Abstract Meaning Representation.
