login

Self-training for biomedical parsing

Published 1 January 2008Open access
David McClosky, Eugene Charniak
Citations115
View PDF

TL;DR

This work self-train the standard Charniak/Johnson Penn-Treebank parser using unlabeled biomedical abstracts, achieving an f-score of 84.3% and a 20% error reduction over the best previous result on biomedical data.

Abstract

Parser self-training is the technique of taking an existing parser, parsing extra data and then creating a second parser by treating the extra data as further training data. Here we apply this technique to parser adaptation. In particular, we self-train the standard Charniak/Johnson Penn-Treebank parser using unlabeled biomedical abstracts. This achieves an f-score of 84.3% on a standard test set of biomedical abstracts from the Genia corpus. This is a 20% error reduction over the best previous result on biomedical data (80.2% on the same test set).

Keywords

Computer Science