login

Learning the Common Structure of Data

Published 30 July 2000
Kristina Lerman, Steven Minton
Citations38

TL;DR

An efficient algorithm is presented that learns structural information about data from positive examples alone and automatically identifies data on Web pages so that the wrapper may be reinduced when the source format changes.

Abstract

The proliferation of online information sources has accentuated the need for tools that automatically validate and recognize data. We present an efficient algorithm that learns structural information about data from positive examples alone. We describe two Web wrapper maintenance applications that employ this algorithm. The first application detects when a wrapper is not extracting correct data. The second application automatically identifies data on Web pages so that the wrapper may be reinduced when the source format changes.

Keywords

Computer ScienceDecision Sciences