login

Estimation of probabilities in the language model of the IBM speech recognition system

IEEE Transactions on Acoustics Speech and Signal ProcessingPublished 1 August 1984
Arthur Nádas
Citations89

TL;DR

The predictive power of the model thus fitted is compared by means of its experimental perplexity to the model as fitted by the Jelinek-Mercer deleted estimator and by the Turing-Good formulas for probabilities of unseen or rarely seen events.

Abstract

The language model probabilities are estimated by an empirical Bayes approach in which a prior distribution for the unknown probabilities is itself estimated through a novel choice of data. The predictive power of the model thus fitted is compared by means of its experimental perplexity [1] to the model as fitted by the Jelinek-Mercer deleted estimator and as fitted by the Turing-Good formulas for probabilities of unseen or rarely seen events.

Keywords

Computer Science