login

Optimizing complex extraction programs over evolving text data

Published 29 June 2009
Fei Chen, Byron J. Gao, AnHai Doan, Jun Yang, Raghu Ramakrishnan
Citations14

TL;DR

Delex is presented, a system that recycles previous IE results to speed up IE over subsequent corpus snapshots, and extensive experiments with both rule-based and learning-based IE programs over two real-world data sets, which demonstrate the utility of the approach.

Abstract

Most information extraction (IE) approaches have considered only static text corpora, over which we apply IE only once. Many real-world text corpora however are dynamic. They evolve over time, and so to keep extracted information up to date we often must apply IE repeatedly, to consecutive corpus snapshots. Applying IE from scratch to each snapshot can take a lot of time. To avoid doing this, we have recently developed Cyclex, a system that recycles previous IE results to speed up IE over subsequent corpus snapshots. Cyclex clearly demonstrated the promise of the recycling idea. The work itself however is limited in that it considers only IE programs that contain a single IE ``blackbox.'' In practice, many IE programs are far more complex, containing multiple IE blackboxes connected in a compositional ``workflow.''

Keywords

Computer Science