login

Selecting articles from the language model training corpus

Published 7 November 2002
Dietrich Klakow
Citations45

TL;DR

A log-likelihood based criterion is suggested to select articles from a training corpus that are suitable to reduce perplexity on a specific task defined by a small target corpus, decreasing the language model size by a factor of 3 at the same time.

Abstract

The paper suggests the use of a log-likelihood based criterion to select articles from a training corpus that are suitable to reduce perplexity on a specific task defined by a small target corpus. This method is not only efficient as an adaptation technique reducing perplexity by 32% and OOV rate from 4.2% to 2.7% but also as a pruning technique, decreasing the language model size by a factor of 3 at the same time.

Keywords

Computer Science