login

Names and similarities on the web

Published 1 January 2006Open access
Marius Paşca, Dekang Lin, Jeffrey P. Bigham, Andrei Lifchits, Alpa Jain
Citations83
View PDF

TL;DR

In a new approach to large-scale extraction of facts from unstructured text, distributional similarities become an integral part of both the iterative acquisition of high-coverage contextual extraction patterns and the validation and ranking of candidate facts.

Abstract

In a new approach to large-scale extraction of facts from unstructured text, distributional similarities become an integral part of both the iterative acquisition of high-coverage contextual extraction patterns, and the validation and ranking of candidate facts. The evaluation measures the quality and coverage of facts extracted from one hundred million Web documents, starting from ten seed facts and using no additional knowledge, lexicons or complex tools.

Keywords

Computer Science