Learning to Denoise Distantly-Labeled Data for Entity Typing
Published 1 January 2019Open access
Yasumasa Onoe, Greg Durrett
Citations68
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
This work proposes a two-stage procedure for handling distantly-labeled data: denoise it with a learned model, then train the final model on clean and denoised distant data with standard supervised training.
Abstract
Yasumasa Onoe, Greg Durrett. Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers). 2019.
Keywords
Computer ScienceDecision Sciences
Deep Residual Learning for Image Recognition
222,082 Citations2016Kaiming He, Xiangyu Zhang +2 more
This work presents a residual learning framework to ease the training of networks that are substantially deeper than those used previously, and provides comprehensive empirical evidence showing that these residual networks are easier to optimize, and can gain accuracy from considerably increased depth.
Neural ComputationLong Short-Term Memory
98,079 Citations1997Sepp Hochreiter, Jürgen Schmidhuber
A novel, efficient, gradient based method called long short-term memory (LSTM) is introduced, which can learn to bridge minimal time lags in excess of 1000 discrete-time steps by enforcing constant error flow through constant error carousels within special units.
UvA-DARE (University of Amsterdam)Adam: A Method for Stochastic Optimization
84,783 Citations2014Diederik P. Kingma, Jimmy Ba
DROPS (Schloss Dagstuhl – Leibniz Center for Informatics)
50,318 Citations2021Mandi, Jayanta, Canoy, Rocsildes +2 more
A simple numeric simulation of DNA-co-polymerized hydrogel shape change and a genetic algorithm that generates and selects large batches of material designs that compete with one another to evolve and converge on optimal objective-matching designs are constructed.
DROPS (Schloss Dagstuhl – Leibniz Center for Informatics)AI-Assisted Pipeline for Dynamic Generation of Trustworthy Health Supplement Content at Scale
45,681 Citations2018Kefallinos, Dionysios, Alexandris, Georgios +6 more
Glove: Global Vectors for Word Representation
33,769 Citations2014Jeffrey Pennington, Richard Socher +1 more
A new global logbilinear regression model that combines the advantages of the two major model families in the literature: global matrix factorization and local context window methods and produces a vector space with meaningful substructure.
32,525 Citations2019Jacob Devlin, Ming‐Wei Chang +2 more
A new language representation model, BERT, designed to pre-train deep bidirectional representations from unlabeled text by jointly conditioning on both left and right context in all layers, which can be fine-tuned with just one additional output layer to create state-of-the-art models for a wide range of tasks.
Extracting and composing robust features with denoising autoencoders
7,343 Citations2008Pascal Vincent, Hugo Larochelle +2 more
This work introduces and motivate a new training principle for unsupervised learning of a representation based on the idea of making the learned representations robust to partial corruption of the input pattern.
Neural NetworksFramewise phoneme classification with bidirectional LSTM and other neural network architectures
5,463 Citations2005Alex Graves, Jürgen Schmidhuber
The main findings are that bidirectional networks outperform unidirectional ones, and Long Short Term Memory (LSTM) is much faster and also more accurate than both standard Recurrent Neural Nets (RNNs) and time-windowed Multilayer Perceptrons (MLPs).
arXiv (Cornell University)Natural Language Processing (almost) from Scratch
5,172 Citations2011Ronan Collobert, Jason Weston +4 more
Natural Language Processing (almost) from Scratch
3,987 Citations2011Ronan Collobert, Jason Weston +4 more
Distant supervision for relation extraction without labeled data
2,913 Citations2009Mike D. Mintz, Steven Bills +2 more
This work investigates an alternative paradigm that does not require labeled corpora, avoiding the domain dependence of ACE-style algorithms, and allowing the use of corpora of any size.
Lecture notes in computer scienceModeling Relations and Their Mentions without Labeled Text
1,331 Citations2010Sebastian Riedel, Limin Yao +1 more
A novel approach to distant supervision that can alleviate the problem of noisy patterns that hurt precision by using a factor graph and applying constraint-driven semi-supervision to train this model without any knowledge about which sentences express the relations in the authors' training KB.
Learning to link with wikipedia
1,266 Citations2008David Milne, Ian H. Witten
This paper explains how machine learning can be used to identify significant terms within unstructured text, and enrich it with links to the appropriate Wikipedia articles, and performs very well, with recall and precision of almost 75%.
IEEE Transactions on Knowledge and Data EngineeringTri-training: exploiting unlabeled data using three classifiers
1,185 Citations2005Zhi‐Hua Zhou, Ming Li
Experiments on UCI data sets and application to the Web page classification task indicate that tri-training can effectively exploit unlabeled data to enhance the learning performance.
Neural Relation Extraction with Selective Attention over Instances
1,097 Citations2016Yankai Lin, Shiqi Shen +3 more
A sentence-level attention-based model for relation extraction that achieves significant and consistent improvements on relation extraction as compared with baselines.
Knowledge-Based Weak Supervision for Information Extraction of Overlapping Relations
923 Citations2011Raphael Hoffmann, Congle Zhang +3 more
A novel approach for multi-instance learning with overlapping relations that combines a sentence-level extraction model with a simple, corpus-level component for aggregating the individual facts is presented.
Proceedings of the VLDB EndowmentSnorkel
761 Citations2017Alexander Ratner, Stephen H. Bach +4 more
Snorkel is a first-of-its-kind system that enables users to train state- of- the-art models without hand labeling any training data and proposes an optimizer for automating tradeoff decisions that gives up to 1.8× speedup per pipeline execution.
Multi-instance Multi-label Learning for Relation Extraction
682 Citations2012Mihai Surdeanu, Julie Tibshirani +2 more
This work proposes a novel approach to multi-instance multi-label learning for RE, which jointly models all the instances of a pair of entities in text and all their labels using a graphical model with latent variables that performs competitively on two difficult domains.
An Improved Non-monotonic Transition System for Dependency Parsing
605 Citations2015Matthew Honnibal, Mark Johnson
A new set of non-monotonic transitions is described that permits a partial parse state to derive a larger set of completed parse trees than previous work, which allows such a parser to escape the “garden paths” that can trap monotonic greedy transition-based dependency parsers.
arXiv (Cornell University)Data Programming: Creating Large Training Sets, Quickly
364 Citations2016Alexander Ratner, Christopher De +3 more
The VLDB JournalSnorkel: rapid training data creation with weak supervision
333 Citations2019Alexander Ratner, Stephen H. Bach +4 more
Ultra-Fine Entity Typing
216 Citations2018Eunsol Choi, Omer Levy +2 more
A model that can predict ultra-fine types is presented, and is trained using a multitask objective that pools the authors' new head-word supervision with prior supervision from entity linking, and achieves state of the art performance on an existing fine-grained entity typing benchmark, and sets baselines for newly-introduced datasets.
Reducing Wrong Labels in Distant Supervision for Relation Extraction
190 Citations2012Shingo Takamatsu, Issei Sato +1 more
A novel generative model is presented that directly models the heuristic labeling process of distant supervision and predicts whether assigned labels are correct or wrong via its hidden variables.
AFET: Automatic Fine-Grained Entity Typing by Hierarchical Partial-Label Embedding
153 Citations2016Xiang Ren, Wenqi He +4 more
This paper proposes a novel embedding method to separately model “clean” and “noisy” mentions, and incorporates the given type hierarchy to induce loss functions.
Proceedings of the AAAI Conference on Artificial IntelligenceTraining Complex Models with Multi-Task Weak Supervision
152 Citations2019Alexander Ratner, Braden Hancock +4 more
This work shows that by solving a matrix completion-style problem, it can recover the accuracies of these multi-task sources given their dependency structure, but without any labeled data, leading to higher-quality supervision for training an end model.
Label Noise Reduction in Entity Typing by Heterogeneous Partial-Label Embedding
150 Citations2016Xiang Ren, Wenqi He +4 more
A global objective is formulated for learning the embeddings from text corpora and knowledge bases, which adopts a novel margin-based loss that is robust to noisy labels and faithfully models type correlation derived from knowledge bases.
Effective Deep Memory Networks for Distant Supervised Relation Extraction
121 Citations2017Xiaocheng Feng, Jiang Guo +3 more
This paper introduces a novel neural approach for distant supervised RE with special focus on attention mechanisms, which includes two major attention-based memory components, which are capable of explicitly capturing the importance of each context word for modeling the representation of the entity pair.
Embedding Methods for Fine Grained Entity Type Classification
118 Citations2015Dani Yogatama, Daniel Gillick +1 more
A new approach based on label embeddings that allows for information sharing among related labels that outperforms state-of-the-art methods on two fine grained entity-classification benchmarks and can exploit the finer-grained labels to improve classification of standard coarse types.
Neural Architectures for Fine-grained Entity Type Classification
110 Citations2017Sonse Shimaoka, Pontus Stenetorp +2 more
This work investigates several neural network architectures for fine-grained entity type classification and establishes that the attention mechanism learns to attend over syntactic heads and the phrase containing the mention, both of which are known to be strong hand-crafted features for this task.
arXiv (Cornell University)Context-Dependent Fine-Grained Entity Type Tagging
107 Citations2014Dan Gillick, Nevena Lazic +3 more
This work proposes the task of context-dependent fine type tagging, where the set of acceptable labels for a mention is restricted to only those deducible from the local context (e.g. sentence or document).
Learning with Noise: Enhance Distantly Supervised Relation Extraction with Dynamic Transition Matrix
101 Citations2017Bingfeng Luo, Yansong Feng +5 more
It is shown that the dynamic transition matrix can effectively characterize the noise in the training data built by distant supervision and can be effectively trained using a novel curriculum learning based method without any direct supervision about the noise.
Hierarchical Losses and New Resources for Fine-grained Entity Typing and Linking
95 Citations2018Shikhar Murty, Patrick Verga +3 more
New methods using real and complex bilinear mappings for integrating hierarchical information are presented, yielding substantial improvement over flat predictions in entity linking and fine-grained entity typing, and achieving new state-of-the-art results for end-to-end models on the benchmark FIGER dataset.
FINET: Context-Aware Fine-Grained Named Entity Typing
90 Citations2015Luciano Del Corro, Abdalghani Abujabal +2 more
FINET generates candidate types using a sequence of multiple extractors, ranging from explicitly mentioned types to implicit types, and subsequently selects the most appropriate using ideas from word-sense disambiguation, and supports the most fine-grained type system so far.
Imposing Label-Relational Inductive Bias for Extremely Fine-Grained Entity Typing
56 Citations2019Wenhan Xiong, Jiawei Wu +5 more
This work introduces a novel label-relational inductive bias, represented by a graph propagation layer that effectively encodes both global label co-occurrence statistics and word-level similarities and shows that a simple modification of the proposed graph layer can improve the performance on a conventional and widely-tested dataset that only includes KB-schema types.
Deep Probabilistic Logic: A Unifying Framework for Indirect Supervision
44 Citations2018Hai Wang, Hoifung Poon
This paper proposes deep probabilistic logic (DPL) as a general framework for indirect supervision, by composing probabilism logic with deep learning, and enables novel combination via infusion of rich domain and linguistic knowledge.
International Conference on Computational LinguisticsCooperative Denoising for Distantly Supervised Relation Extraction
32 Citations2018Kai Lei, Daoyuan Chen +5 more
CORD, a novelCOopeRativeDenoising framework, which consists two base networks leveraging text corpus and knowledge graph respectively, and a cooperative module involving their mutual learning by the adaptive bi-directional knowledge distillation and dynamic ensemble with noisy-varying instances achieves substantial improvement over state-of-the-art methods.
arXiv (Cornell University)Denoising Distant Supervision for Relation Extraction via Instance-Level Adversarial Training
17 Citations2018Xu Han, Zhiyuan Liu +1 more
This paper proposes a novel adversarial training mechanism over instances for relation extraction to alleviate the noise issue and develops a denoising method that can better discriminate those informative instances from noisy ones.
Fine-Grained Entity Typing with High-Multiplicity Assignments
14 Citations2017Maxim Rabinovich, Dan Klein
A set-prediction approach is introduced to the high-multiplicity regime inherent in data sources such as Wikipedia that have semi-open type systems and it is shown that this model outperforms unstructured baselines on a new Wikipedia-based fine-grained typing corpus.
