Improving Named Entity Recognition by External Context Retrieving and Cooperative Learning
Published 1 January 2021Open access
Xinyu Wang, Yong Jiang, Nguyễn Bách, Tao Wang, Zhongqiang Huang, Fei Huang
Citations132
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
This paper finds empirically that the contextual representations computed on the retrieval-based input view, constructed through the concatenation of a sentence and its external contexts, can achieve significantly improved performance compared to the original input view based only on the sentence.
Abstract
Xinyu Wang, Yong Jiang, Nguyen Bach, Tao Wang, Zhongqiang Huang, Fei Huang, Kewei Tu. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Keywords
Computer Science
32,525 Citations2019Jacob Devlin, Ming‐Wei Chang +2 more
A new language representation model, BERT, designed to pre-train deep bidirectional representations from unlabeled text by jointly conditioning on both left and right context in all layers, which can be fine-tuned with just one additional output layer to create state-of-the-art models for a wide range of tasks.
DROPS (Schloss Dagstuhl – Leibniz Center for Informatics)HISTORIAE, History of Socio-Cultural Transformation as Linguistic Data Science. A Humanities Use Case
17,334 Citations2019Yinhan Liu, Myle Ott +8 more
This work considers the task of building machine learning models to automatically select the best combination for a problem instance and contributes to the automatic learning of instance features directly from the high-level representation of a problem instance using a transformer encoder.
arXiv (Cornell University)Distilling the Knowledge in a Neural Network
13,965 Citations2015Geoffrey E. Hinton, Oriol Vinyals +1 more
This work shows that it can significantly improve the acoustic model of a heavily used commercial system by distilling the knowledge in an ensemble of models into a single model and introduces a new type of ensemble composed of one or more full models and many specialist models which learn to distinguish fine-grained classes that the full models confuse.
ScholarlyCommons (University of Pennsylvania)Conditional Random Fields: Probabilistic Models for Segmenting and Labeling Sequence Data
12,978 Citations2001John Lafferty, Andrew McCallum +1 more
This work presents iterative parameter estimation algorithms for conditional random fields and compares the performance of the resulting models to HMMs and MEMMs on synthetic and natural-language data.
arXiv (Cornell University)Decoupled Weight Decay Regularization
9,113 Citations2017Ilya Loshchilov, Frank Hutter
This work proposes a simple modification to recover the original formulation of weight decay regularization by decoupling the weight decay from the optimization steps taken w.r.t. the loss function, and provides empirical evidence that this modification substantially improves Adam's generalization performance.
BioinformaticsBioBERT: a pre-trained biomedical language representation model for biomedical text mining
7,029 Citations2019Jinhyuk Lee, Wonjin Yoon +5 more
This article introduces BioBERT (Bidirectional Encoder Representations from Transformers for Biomedical Text Mining), which is a domain-specific language representation model pre-trained on large-scale biomedical corpora that largely outperforms BERT and previous state-of-the-art models in a variety of biomedical text mining tasks when pre- trained on biomedical Corpora.
Combining labeled and unlabeled data with co-training
5,604 Citations1998Avrim Blum, Tom M. Mitchell
Neural Architectures for Named Entity Recognition
4,428 Citations2016Guillaume Lample, Miguel Ballesteros +3 more
Comunicacio presentada a la 2016 Conference of the North American Chapter of the Association for Computational Linguistics, celebrada a San Diego (CA, EUA) els dies 12 a 17 of juny 2016.
arXiv (Cornell University)Affordance-Compiled Intelligence: Observable-Only Cognitive Impedance Matching for No-Meta LLM-Integrated Systems
2,995 Citations2026Patrick Lewis, Ethan Perez +10 more
End-to-end Sequence Labeling via Bi-directional LSTM-CNNs-CRF
2,589 Citations2016Xuezhe Ma, Eduard Hovy
A novel neutral network architecture is introduced that benefits from both word- and character-level representations automatically, by using combination of bidirectional LSTM, CNN and CRF, thus making it applicable to a wide range of sequence labeling tasks.
Model compression
2,087 Citations2006Cristian Buciluǎ, Rich Caruana +1 more
This work presents a method for "compressing" large, complex ensembles into smaller, faster models, usually without significant loss in performance.
arXiv (Cornell University)BERTScore: Evaluating Text Generation with BERT
2,045 Citations2019Tianyi Zhang, Varsha Kishore +3 more
arXiv (Cornell University)Introduction to the CoNLL-2002 Shared Task: Language-Independent Named Entity Recognition
1,574 Citations2002Erik F. Tjong Kim Sang
International Conference on Computational LinguisticsContextual String Embeddings for Sequence Labeling
1,004 Citations2018Alan Akbik, Duncan A. J. Blythe +1 more
This paper proposes to leverage the internal states of a trained character language model to produce a novel type of word embedding which they refer to as contextual string embeddings, which are fundamentally model words as sequences of characters and are contextualized by their surrounding text.
arXiv (Cornell University)A Survey on Multi-view Learning
999 Citations2013Chang Xu, Dacheng Tao +1 more
By exploring the consistency and complementary properties of different views, multi-View learning is rendered more effective, more promising, and has better generalization ability than single-view learning.
Neural Computing and ApplicationsA survey of multi-view machine learning
852 Citations2013Shiliang Sun
This paper reviews theories developed to understand the properties and behaviors of multi-view learning and gives a taxonomy of approaches according to the machine learning mechanisms involved and the fashions in which multiple views are exploited.
DatabaseBioCreative V CDR task corpus: a resource for chemical disease relation extraction
839 Citations2016Jiao Li, Yueping Sun +8 more
The BC5CDR corpus was successfully used for the BioCreative V challenge tasks and should serve as a valuable resource for the text-mining research community.
Journal of Biomedical InformaticsNCBI disease corpus: A resource for disease name recognition and concept normalization
804 Citations2014Rezarta Islamaj, Robert Leaman +1 more
The results show that the NCBI disease corpus has the potential to significantly improve the state-of-the-art in disease name recognition and normalization research, by providing a high-quality gold standard thus enabling the development of machine-learning based approaches for such tasks.
arXiv (Cornell University)BERTScore: Evaluating Text Generation with BERT
603 Citations2020Tianyi Zhang, Varsha Kishore +3 more
This work proposes BERTScore, an automatic evaluation metric for text generation that correlates better with human judgments and provides stronger model selection performance than existing metrics.
Dice Loss for Data-imbalanced NLP Tasks
557 Citations2020Xiaoya Li, Xiaofei Sun +4 more
This paper proposes to use dice loss in replacement of the standard cross-entropy objective for data-imbalanced NLP tasks, based on the Sørensen--Dice coefficient or Tversky index, which attaches similar importance to false positives and false negatives, and is more immune to the data-IMbalance issue.
LUKE: Deep Contextualized Entity Representations with Entity-aware Self-attention
551 Citations2020Ikuya Yamada, Akari Asai +3 more
New pretrained contextualized representations of words and entities based on the bidirectional transformer, and an entity-aware self-attention mechanism that considers the types of tokens (words or entities) when computing attention scores are proposed.
Named Entity Recognition as Dependency Parsing
425 Citations2020Juntao Yu, Bernd Bohnet +1 more
Ideas from graph-based dependency parsing are used to provide the model a global view on the input via a biaffine model and show that the model works well for both nested and flat NER, through evaluation on 8 corpora and achieving SoTA performance on all of them.
Semi-Supervised Sequence Modeling with Cross-View Training
393 Citations2018Kevin B. Clark, Minh-Thang Luong +2 more
Cross-View Training (CVT), a semi-supervised learning algorithm that improves the representations of a Bi-LSTM sentence encoder using a mix of labeled and unlabeled data, is proposed and evaluated, achieving state-of-the-art results.
Results of the WNUT2017 Shared Task on Novel and Emerging Entity Recognition
366 Citations2017Leon Derczynski, Eric Nichols +2 more
The goal of this task is to provide a definition of emerging and of rare entities, and based on that, also datasets for detecting these entities and to evaluate the ability of participating entries to detect and classify novel and emerging named entities in noisy text.
Pooled Contextualized Embeddings for Named Entity Recognition
307 Citations2019Alan Akbik, Tanja Bergmann +1 more
This work proposes a method in which it dynamically aggregate contextualized embeddings of each unique string that the authors encounter and uses a pooling operation to distill a ”global” word representation from all contextualized instances.
Retrieve, Rerank and Rewrite: Soft Template Based Neural Summarization
224 Citations2018Ziqiang Cao, Wenjie Li +2 more
This paper uses a popular IR platform to use existing summaries as soft templates to guide the seq2seq model, and extends the framework to jointly conduct template Reranking and template-aware summary generation (Rewriting).
Retrieve and Refine: Improved Sequence Generation Models For Dialogue
176 Citations2018Jason Weston, Emily Dinan +1 more
This work develops a model that combines the two approaches to avoid both their deficiencies: first retrieve a response and then refine it – the final sequence generator treating the retrieval as additional context.
Proceedings of the AAAI Conference on Artificial IntelligenceSearch Engine Guided Neural Machine Translation
152 Citations2018Jiatao Gu, Yong Wang +2 more
An attention-based neural machine translation model is extended by allowing it to access an entire training set of parallel sentence pairs even after training, and significantly outperforms the baseline approach.
Automated Concatenation of Embeddings for Structured Prediction
134 Citations2021Xinyu Wang, Yong Jiang +5 more
This paper proposes Automated Concatenation of Embeddings (ACE) to automate the process of finding better concatenations of embeddings for structured prediction tasks, based on a formulation inspired by recent progress on neural architecture search.
Cross-Domain NER using Cross-Domain Language Modeling
121 Citations2019Jia Chen, Xiaobo Liang +1 more
This work considers using cross- domain LM as a bridge cross-domains for NER domain adaptation, performing cross-domain and cross-task knowledge transfer by designing a novel parameter generation network and shows that this method can effectively extract domain differences from cross- domains LM contrast, allowing unsupervised domain adaptation while also giving state-of-the-art results among supervised domain adaptation methods.
Dual Adversarial Neural Transfer for Low-Resource Named Entity Recognition
116 Citations2019Joey Tianyi Zhou, Hao Zhang +5 more
Two variants of DATNet are investigated to explore effective feature fusion between high and low resource, and a novel Generalized Resource-Adversarial Discriminator (GRAD) is proposed to address the noisy and imbalanced training data.
Guiding Neural Machine Translation with Retrieved Translation Pieces
116 Citations2018Jingyi Zhang, Masao Utiyama +3 more
This paper proposes a simple, fast, and effective method for recalling previously seen translation examples and incorporating them into the NMT decoding process, and compares favorably to another alternative retrieval-based method with respect to accuracy, speed, and simplicity of implementation.
Digital Access to LibrariesResults of the WNUT16 Named Entity Recognition Shared Task
112 Citations2016Benjamin Strauss, Bethany Toma +3 more
The shared task, annotation process and dataset statistics are outlined, and a high-level overview of the participating systems for each shared task is provided.
arXiv (Cornell University)A Retrieve-and-Edit Framework for Predicting Structured Outputs
102 Citations2018Tatsunori Hashimoto, Kelvin Guu +2 more
This work proposes an approach that first retrieves a training example based on the input and then edits it to the desired output, and shows that on a new autocomplete task for GitHub Python code and the Hearthstone cards benchmark, retrieve-and-edit significantly boosts the performance of a vanilla sequence-to-sequence model on both tasks.
CrossWeigh: Training Named Entity Tagger from Imperfect Annotations
99 Citations2019Zihan Wang, Jingbo Shang +4 more
This study dives deep into one of the widely-adopted NER benchmark datasets, CoNLL03 NER, and proposes a simple yet effective framework, CrossWeigh, to handle label mistakes during NER model training.
Retrieval-Based Neural Code Generation
91 Citations2018Shirley Anugrah Hayati, Raphaël Olivier +4 more
Recode, a method based on subtree retrieval that makes it possible to explicitly reference existing code examples within a neural code generation model, is introduced.
Named Entity Recognition for Social Media Texts with Semantic Augmentation
69 Citations2020Yuyang Nie, Yuanhe Tian +3 more
A neural-based approach to NER for social media texts where both local and augmented semantics are taken into account, and an attentive semantic augmentation module and a gate module to encode and aggregate such information are proposed.
Boosting Neural Machine Translation with Similar Translations
59 Citations2020Jitao XU, Josep Crego +1 more
This paper explores data augmentation methods for training Neural Machine Translation to make use of similar translations, in a comparable way a human translator employs fuzzy matches, and shows that translations based on fuzzy matching provide the model with “copy” information while translationsbased on embedding similarities tend to extend the translation “context”.
Retrieval-guided Dialogue Response Generation via a Matching-to-Generation Framework
58 Citations2019Deng Cai, Yan Wang +4 more
A novel framework in which the skeleton extraction is made by an interpretable matching model and the following skeleton-guided response generation is accomplished by a separately trained generator is presented.
Bilingual Dictionary Based Neural Machine Translation without Using Parallel Sentences
50 Citations2020Jitao Xu, Josep-Maria Crego +6 more
Coupling Retrieval and Meta-Learning for Context-Dependent Semantic Parsing
44 Citations2019Daya Guo, Duyu Tang +3 more
An approach to incorporate retrieved datapoints as supporting evidence for context-dependent semantic parsing, such as generating source code conditioned on the class environment, and shows that both the context-aware retriever and the meta-learning strategy improve accuracy.
Reinforcement-based denoising of distantly supervised NER with partial annotation
32 Citations2019Farhad Nooralahzadeh, Jan Tore Lønning +1 more
This paper adopts a technique of partial annotation to address false negative cases and implements a reinforcement learning strategy with a neural network policy to identify false positive instances and establishes a new state-of-the-art on four benchmark datasets taken from different domains and different languages.
Structure-Level Knowledge Distillation For Multilingual Sequence Labeling
29 Citations2020Xinyu Wang, Yong Jiang +4 more
This paper proposes two novel KD methods based on structure-level information that approximately minimizes the distance between the student’s and the teachers’ structure- level probability distributions, and aggregates theructure-level knowledge to local distributions and minimizesThe distance between two local probability distributions.
Retrieval-Augmented Controllable Review Generation
19 Citations2020Jihyeok Kim, Seungtaek Choi +2 more
This paper proposes to additionally leverage references, which are selected from a large pool of texts labeled with one of the attributes, as textual information that enriches inductive biases of given attributes.
arXiv (Cornell University)BioFLAIR: Pretrained Pooled Contextualized Embeddings for Biomedical Sequence Labeling Tasks
16 Citations2019Shreyas Sharma, Ron Daniel
It is found that with the provided embeddings, FLAIR performs on-par with the BERT networks - even establishing a new state of the art on one benchmark.
arXiv (Cornell University)BioFLAIR: Pretrained Pooled Contextualized Embeddings for Biomedical\n Sequence Labeling Tasks
14 Citations2019Shreyas Sharma, Ron W. Daniel
Multi-View Cross-Lingual Structured Prediction with Minimum Supervision
8 Citations2021Zechuan Hu, Yong Jiang +5 more
This paper proposes a multi-view framework, by leveraging a small number of labeled target sentences, to effectively combine multiple source models into an aggregated source view at different granularity levels (language, sentence, or sub-structure), and transfer it to a target view based on a task-specific model.
Structural Knowledge Distillation: Tractably Distilling Information for Structured Predictor
8 Citations2021Xinyu Wang, Yong Jiang +7 more
A factorized form of the knowledge distillation objective for structured prediction, which is tractable for many typical choices of the teacher and student models and shows the tractability and empirical effectiveness of structural knowledgedistillation between sequence labeling and dependency parsing models.
arXiv (Cornell University)Structural Knowledge Distillation: Tractably Distilling Information for Structured Predictor
6 Citations2020Xinyu Wang, Yong Jiang +7 more
