Weakly-supervised acquisition of labeled class instances using graph random walks
Published 1 January 2008Open access
Partha Talukdar, Joseph Reisinger, Marius Paşca, Deepak Ravichandran, Rahul Bhagat, Fernando Pereira
Citations133
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
This acquisition method significantly improves coverage compared to a previous set of labeled classes and instances derived from free text, while achieving comparable precision.
Abstract
We present a graph-based semi-supervised label propagation algorithm for acquiring open-domain labeled classes and their instances from a combination of unstructured and structured text sources. This acquisition method significantly improves coverage compared to a previous set of labeled classes and instances derived from free text, while achieving comparable precision.
Keywords
Computer Science
Semi-supervised learning using Gaussian fields and harmonic functions
3,377 Citations2003Xiaojin Zhu, Zoubin Ghahramani +1 more
Automatic acquisition of hyponyms from large text corpora
3,283 Citations1992Marti A. Hearst
A set of lexico-syntactic patterns that are easily recognizable, that occur frequently and across text genre boundaries, and that indisputably indicate the lexical relation of interest are identified.
Information Processing & ManagementReal life, real users, and real needs: a study and analysis of user queries on the web
1,318 Citations2000Bernard J. Jansen, Amanda Spink +1 more
A failure analysis was conducted, identifying trends among user mistakes, and a summary of findings and a discussion of the implications of these findings were concluded.
Learning dictionaries for information extraction by multi-level bootstrapping
688 Citations1999Ellen Riloff, Rosie Jones
A multilevel bootstrapping algorithm is presented that generates both the semantic lexicon and extraction patterns simultaneously simultaneously and produces high-quality dictionaries for several semantic categories.
Proceedings of the VLDB EndowmentWebTables
634 Citations2008Michael Cafarella, Alon Halevy +3 more
The WEBTABLES system develops new techniques for keyword search over a corpus of tables, and shows that they can achieve substantially higher relevance than solutions based on a traditional search engine.
Partially labeled classification with Markov random walks
563 Citations2001Martin Szummer, Tommi Jaakkola
This work combines a limited number of labeled examples with a Markov random walk representation over the unlabeled examples and develops and compares several estimation criteria/algorithms suited to this representation.
Video suggestion and discovery for youtube
444 Citations2008Shumeet Baluja, Rohan Seth +6 more
A novel method based upon the analysis of the entire user-video graph to provide personalized video suggestions for users, termed Adsorption, provides a simple method to efficiently propagate preference information through a variety of graphs.
Language-Independent Set Expansion of Named Entities Using the Web
207 Citations2007Richard C. Wang, William W. Cohen
This paper proposes a novel method for expanding sets of named entities that can be applied to semi-structured documents written in any markup language and in any human language and shows that this system is superior to Google Sets in terms of mean average precision.
Concept discovery from text
185 Citations2002Dekang Lin, Patrick Pantel
A clustering algorithm called CBC (Clustering By Committee) that automatically discovers concepts from text that outperforms several well-known clustering algorithms in cluster quality.
Low-Distortion Embeddings of Finite Metric Spaces
120 Citations2004Piotr Indyk, Jiřı́ Matoušek
International Joint Conference on Artificial IntelligenceUsing decision trees for conference resolution
97 Citations1995Joseph F. McCarthy, Wendy G. Lehnert
arXiv (Cornell University)Using Decision Trees for Coreference Resolution
93 Citations1995Joseph F. McCarthy, Wendy G. Lehnert
RESOLVE is described, a system that uses decision trees to learn how to classify coreferent phrases in the domain of business joint ventures and provides a framework that facilitates the exploration of the types of knowledge that are useful for solving the coreference problem.
The rendezvous algorithm
90 Citations2007Arik Azran
A new approach for estimating a distribution over the missing labels where data points are viewed as nodes of a graph, and pairwise similarities are used to derive a transition probability matrix P for a Markov random walk between them.
Using corpus-derived name lists for named entity recognition
41 Citations2000Mark Stevenson, Robert Gaizauskas
This paper describes experiments to establish the performance of a named entity recognition system which builds categorized lists of names from manually annotated training data and shows that by using simple filtering techniques for improving the automatically acquired lists, substantial performance benefits can be achieved.
Finding cars, goddesses and enzymes: parametrizable acquisition of labeled instances for open-domain information extraction
40 Citations2008Benjamin Van Durme, Marius Paşca
Experimental results show the procedure allows for a parametric trade-off between high precision and expanded recall.
CORE Scholar (Wright State University)Lightly-Supervised Attribute Extraction
27 Citations2007Kedar Bellare, Partha Talukdar +5 more
This paper introduces lightly-supervised methods for extracting entity attributes from natural language text using a bootstrapping approach, and is able to extract large numbers of attributes of different entities at fairly high precision from a large natural language corpus.
