Adaptive language modeling using minimum discriminant estimation
Published 1 January 1992Open access
S. Della Pietra, V. Della Pietra, R. L. Mercer, Salim Roukos
Citations79
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
Abstract
We present an algorithm to adapt a n-gram language model to a document as it is dictated. The observed partial document is used to estimate a unigram distribution for the words that already occurred. Then, we find the closest n-gram distribution to the static n-gram distribution (using the discrimination information distance measure) and that satisfies the marginal constraints derived from the document. The resulting minimum discrimination information model results in a perplexity of 208 instead of 290 for the static trigram model on a document of 321 words.
Keywords
Computer Science
The Annals of Mathematical StatisticsGeneralized Iterative Scaling for Log-Linear Models
1,203 Citations1972J. N. Darroch, D. Ratcliff
A dynamic language model for speech recognition
125 Citations1991F. Jelinek, Bernard Mérialdo +2 more
This model is called a cache trigram language model (CTLM) since it is caching the recent history of words and it is found that the CTLM reduces the perplexity of a dictated document by 23%.
Speech recognition and the frequency of recently used words
46 Citations1988Roland Kühn
A modification of the Markov approach, which assigns higher probabilities to recently used words, is proposed and tested against a pure Markov model.
Probabilistic models of short and long distance word dependencies in running text
26 Citations1989Julien Kupiec
Two complementary models that represent dependencies between words in local and non-local contexts are described, which are useful for doing prediction in systems using large vocabularies.
