A Bayesian Interpretation of Interpolated Kneser-Ney
Published 1 January 2006
YW Teh
Citations108
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
Abstract
tutorial article, which has been submitted for publication in a journal or for consideration by the commissioning organization. The report represents the ideas of its author, and should not be taken as the official views of the School or the University. Any discussion of the content of the report should be sent to the author, at the address shown on the cover.
Keywords
Computer ScienceMathematics
Bayesian Data Analysis
13,688 Citations2003Andrew Gelman, John B. Carlin +2 more
Journal of the American Statistical AssociationHierarchical Dirichlet Processes
3,567 Citations2006Yee Whye Teh, Michael I. Jordan +2 more
This work considers problems involving groups of data where each observation within a group is a draw from a mixture model and where it is desirable to share mixture components between groups, and considers a hierarchical model, specifically one in which the base measure for the childDirichlet processes is itself distributed according to a Dirichlet process.
Computational LinguisticsA maximum entropy approach to natural language processing
3,120 Citations1996Adam Berger, Vincent J. Della Pietra +1 more
A maximum-likelihood approach for automatically constructing maximum entropy models is presented and how to implement this approach efficiently is described, using as examples several problems in natural language processing.
Computer Speech & LanguageAn empirical study of smoothing techniques for language modeling
2,077 Citations1999Stanley F. Chen, Joshua Goodman
Applied Physics Letters10.1162/153244303322533223
1,771 Citations2000
This work proposes to fight the curse of dimensionality by learning a distributed representation for words which allows each training sentence to inform the model about an exponential number of semantically neighboring sentences.
Journal of the American Statistical AssociationGibbs Sampling Methods for Stick-Breaking Priors
1,556 Citations2001Hemant Ishwaran, Lancelot F. James
Two general types of Gibbs samplers that can be used to fit posteriors of Bayesian hierarchical models based on stick-breaking priors are presented and the blocked Gibbs sampler, based on an entirely different approach that works by directly sampling values from the posterior of the random measure.
Improved backing-off for M-gram language modeling
1,490 Citations2002Reinhard Kneser, Hermann Ney
This paper proposes to use distributions which are especially optimized for the task of back-off, which are quite different from the probability distributions that are usually used for backing-off.
Maximum Entropy Markov Models for Information Extraction and Segmentation
1,333 Citations2000Andrew McCallum, Dayne Freitag +1 more
A new Markovian sequence model is presented that allows observations to be represented as arbitrary overlapping features (such as word, capitalization, formatting, part-of-speech), and defines the conditional probability of state sequences given observation sequences.
The Annals of ProbabilityThe two-parameter Poisson-Dirichlet distribution derived from a stable subordinator
1,255 Citations1997Jim Pitman, Marc Yor
Lecture notes in mathematicsCombinatorial Stochastic Processes
954 Citations2006Jim Pitman, Jean Picard
Proceedings of the IEEETwo decades of statistical language modeling: where do we go from here?
726 Citations2000Roni Rosenfeld
A Bayesian approach to integration of linguistic theories with data is argued for inStatistical language models estimate the distribution of various natural language phenomena for the purpose of speech recognition and other language technologies.
Computer Speech & LanguageA bit of progress in language modeling
463 Citations2001Joshua Goodman
A combination of all techniques together to a Katz smoothed trigram model with no count cutoffs achieves perplexity reductions between 38 and 50% (1 bit of entropy), depending on training data size, as well as a word error rate reduction of 8.9%.
Immediate-head parsing for language models
299 Citations2001Eugene Charniak
It is suggested that improvement of the underlying parser should significantly improve the model's perplexity and that even in the near term there is a lot of potential for improvement in immediate-head language models.
Factored language models and generalized parallel backoff
283 Citations2003Jeff Bilmes, Katrin Kirchhoff
Initial perplexity results on both CallHome Arabic and on Penn Treebank Wall Street Journal articles are provided, and FLMs with GPB can produce bigrams with significantly lower perplexity, sometimes lower than highly-optimized baseline trigrams.
Research Showcase @ Carnegie Mellon University (Carnegie Mellon University)Adaptive Statistical Language Modeling: A Maximum Entropy Approach
276 Citations2018Roni Rosenfeld
This thesis views language as an information source which emits a stream of symbols from a finite alphabet (the vocabulary), and applies the principle of Maximum Entropy to identify and exploit sources of information in the language stream, so as to minimize its perceived entropy.
Computer Speech & LanguageA comparison of the enhanced Good-Turing and deleted estimation methods for estimating probabilities of English bigrams
240 Citations1991Kenneth Church, William A. Gale
The enhanced Good-Turing method is introduced, which is three to four times as efficient in its use of data as the enhanced deleted estimate method and provides accurate predictions of the variances of the standard probabilities estimated from the test corpus.
IEEE Transactions on Speech and Audio ProcessingA survey of smoothing techniques for ME models
199 Citations2000S.F. Chen, Roni Rosenfeld
Over a large number of data sets, it is found that fuzzy ME smoothing performs as well as or better than all other algorithms under consideration.
Neural ComputationStructure Learning in Conditional Probability Models via an Entropic Prior and Parameter Extinction
159 Citations1999Matthew Brand
An entropic prior is introduced for multinomial parameter estimation problems and the resulting models show superior generalization to held-out test data, and a guarantee that any such deletion will increase the posterior probability of the model.
