login

Extracting Patterns and Relations from the World Wide Web

Lecture notes in computer sciencePublished 1 January 1999
Sergey Brin
Citations1,007
SJR quartileQ2
SJR score0.35
SNIP0.55

TL;DR

This paper presents a technique which exploits the duality between sets of patterns and relations to grow the target relation starting from a small sample and uses it to extract a relation of (author,title) pairs from the World Wide Web.

Abstract

The World Wide Web is a vast resource for information. At the same time it is extremely distributed. A particular type of data such as restaurant lists may be scattered across thousands of independent information sources in many different formats. In this paper, we consider the problem of extracting a relation for such a data type from all of these sources automatically. We present a technique which exploits the duality between sets of patterns and relations to grow the target relation starting from a small sample. To test our technique we use it to extract a relation of (author,title) pairs from the World Wide Web.

Keywords

Computer Science