login

Bootstrapping statistical parsers from small datasets

Published 1 January 2003Open access
Mark Steedman, Miles Osborne, Anoop Sarkar, Stephen Clark, Rebecca Hwa, Julia Hockenmaier
Citations153
View PDF

TL;DR

Experimental results show that unlabelled sentences can be used to improve the performance of statistical parsers and it is shown that boot-strapping continues to be useful, even though no manually produced parses from the target domain are used.

Abstract

We present a practical co-training method for bootstrapping statistical parsers using a small amount of manually parsed training material and a much larger pool of raw sentences. Experimental results show that unlabelled sentences can be used to improve the performance of statistical parsers. In addition, we consider the problem of boot-strapping parsers when the manually parsed training material is in a different domain to either the raw sentences or the testing material. We show that boot-strapping continues to be useful, even though no manually produced parses from the target domain are used.

Keywords

Computer Science