A Bit of Progress in Language Modeling Extended Version
Seminars in Clinical NeuropsychiatryPublished 1 August 2001
Joshua Goodman
Citations114
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
Abstract
1.1 Overview Language modeling is the art of determining the probability of a sequence of words. This is useful in a large variety of areas including speech recognition,
Keywords
Computer Science
Numerical recipes in C
15,253 Citations1994William H. Press, Saul A. Teukolsky +2 more
The Diskette v 2.06, 3.5''[1.44M] for IBM PC, PS/2 and compatibles [DOS] Reference Record created on 2004-09-07, modified on 2016-08-08.
BiometrikaTHE POPULATION FREQUENCIES OF SPECIES AND THE ESTIMATION OF POPULATION PARAMETERS
3,227 Citations1953I. J. Good
Class-based n -gram models of natural language
2,893 Citations1992Peter F. Brown, P.V. deSouza +3 more
This work addresses the problem of predicting a word from previous words in a sample of text and discusses n-gram models based on classes of words, finding that these models are able to extract classes that have the flavor of either syntactically based groupings or semanticallybased groupings, depending on the nature of the underlying statistics.
Computer Speech & LanguageAn empirical study of smoothing techniques for language modeling
2,077 Citations1999Stanley F. Chen, Joshua Goodman
A statistical approach to machine translation
1,699 Citations1990Peter F. Brown, John Cocke +6 more
The application of the statistical approach to translation from French to English and preliminary results are described and the results are given.
IEEE Transactions on Acoustics Speech and Signal ProcessingEstimation of probabilities from sparse data for the language model component of a speech recognizer
1,645 Citations1987Slava M. Katz
The model offers, via a nonlinear recursive procedure, a computation and space efficient solution to the problem of estimating probabilities from sparse data, and compares favorably to other proposed methods.
Improved backing-off for M-gram language modeling
1,490 Citations2002Reinhard Kneser, Hermann Ney
This paper proposes to use distributions which are especially optimized for the task of back-off, which are quite different from the probability distributions that are usually used for backing-off.
The Annals of Mathematical StatisticsGeneralized Iterative Scaling for Log-Linear Models
1,203 Citations1972J. N. Darroch, D. Ratcliff
A Neural Probabilistic Language Model
1,156 Citations2000Yoshua Bengio, Réjean Ducharme +1 more
Distributional clustering of English words
994 Citations1993Fernando Pereira, Naftali Tishby +1 more
A stochastic parts program and noun phrase parser for unrestricted text
974 Citations1988Kenneth Church
A program that tags each word in an input sentence with the most likely part of speech has been written and performance is encouraging; a 400-word sample is presented and is judged to be 99.5% correct.
Proceedings of the IEEETwo decades of statistical language modeling: where do we go from here?
726 Citations2000Roni Rosenfeld
A Bayesian approach to integration of linguistic theories with data is argued for inStatistical language models estimate the distribution of various natural language phenomena for the purpose of speech recognition and other language technologies.
Computer Speech & LanguageThe SPHINX-II speech recognition system: an overview
370 Citations1993Xuedong Huang, F. Alleva +4 more
The SPHINX-II speech recognition system is reviewed and recent efforts on improved speech recognition are summarized.
Proceedings of the IEEEExploiting latent semantic information in statistical language modeling
352 Citations2000J.R. Bellegarda
This paper focuses on the use of latent semantic analysis, a paradigm that automatically uncovers the salient semantic relationships between words and documents in a given corpus, and proposes an integrative formulation for harnessing this synergy.
A Gaussian Prior for Smoothing Maximum Entropy Models
347 Citations1999Stanley F. Chen, Ronald Rosenfeld
Over a large number of data sets, it is found that an ME smoothing method proposed to us by Lafferty performs as well as or better than all other algorithms under consideration.
Immediate-head parsing for language models
299 Citations2001Eugene Charniak
It is suggested that improvement of the underlying parser should significantly improve the model's perplexity and that even in the near term there is a lot of potential for improvement in immediate-head language models.
Research Showcase @ Carnegie Mellon University (Carnegie Mellon University)Adaptive Statistical Language Modeling: A Maximum Entropy Approach
276 Citations2018Roni Rosenfeld
This thesis views language as an information source which emits a stream of symbols from a finite alphabet (the vocabulary), and applies the principle of Maximum Entropy to identify and exploit sources of information in the language stream, so as to minimize its perceived entropy.
Natural Language EngineeringA hierarchical Dirichlet language model
276 Citations1995David Mackay, Linda C. Bauman Peto
A hierarchical probabilistic model whose predictions are similar to those of the popular language modelling procedure known as ‘smoothing’ is discussed, and the methods prove to be about equally accurate, with the hierarchical model using fewer computational resources.
arXiv (Cornell University)Entropy-based Pruning of Backoff Language Models
267 Citations2000Andreas Stolcke
It is shown that the relative entropy resulting from pruning a single N-gram can be computed exactly and efficiently for backoff models and shown that a production-quality Hub4 LM can be reduced to 26% its original size without increasing recognition error.
A spelling correction program based on a noisy channel model
265 Citations1990Mark D. Kernighan, Kenneth Church +1 more
A new program is described, correct, which takes words rejected by the Unix® spell program, proposes a list of candidate corrections, and sorts them by probability, and the probability scores are the novel contribution of this work.
Classes for fast maximum entropy training
171 Citations2002J.T. Goodman
This paper presents a speedup technique: changing the form of the model to use classes, which leads to fewer nonzero indicator functions, and faster normalization, achieving speedups of up to a factor of 35 over one of the best previous techniques.
Exploiting Syntactic Structure for Language Modeling
164 Citations1998Ciprian Chelba, Frederick Jelinek
IEEE Transactions on Speech and Audio ProcessingModeling long distance dependence in language: topic mixtures versus dynamic cache models
158 Citations1999Rajah Iyer, Mari Ostendorf
A topic-dependent, sentence-level mixture language model which takes advantage of the topic constraints in a sentence or article, and introduces topic-dependent dynamic adaptation techniques in the framework of the mixture model, using n-gram caches and content word unigram caches.
ArXiv.orgAggregate and mixed-order Markov models for statistical language processing
139 Citations1997Lawrence K. Saul, Fernando Pereira
This work considers the use of language models whose size and accuracy are intermediate between different order n-gram models and examines smoothing procedures in which these models are interposed between different orders.
A novel word clustering algorithm based on latent semantic analysis
101 Citations2002J.R. Bellegarda, John Butzberger +3 more
A new approach is proposed for the clustering of words in a given vocabulary based on a paradigm first formulated in the context of information retrieval, called latent semantic analysis, which leads to a parsimonious vector representation of each word in a suitable vector space.
Using a stochastic context-free grammar as a language model for speech recognition
93 Citations2002Daniel Jurafsky, Chuck Wooters +5 more
An algorithm for using a probabilistic Earley parser and a stochastic context-free grammar (SCFG) to generate word transition probabilities at each frame for a Viterbi decoder and it is shown that using an SCFG as a language model improves the word error rate.
Computer Speech & LanguageWhole-sentence exponential language models: a vehicle for linguistic-statistical integration
93 Citations2001Ronald Rosenfeld, Stanley F. Chen +1 more
An exponential language model which models a whole sentence or utterance as a single unit is introduced, and a novel procedure for feature selection is presented, which exploits discrepancies between the existing model and the training corpus.
IEEE Transactions on Speech and Audio ProcessingVariable n-grams and extensions for conversational speech language modeling
89 Citations2000Man-Hung Siu, Mari Ostendorf
Experiments show that using the extended variable n-gram design algorithm to conversational speech results in a language model that captures 4-gram context with less than half the parameters of a standard trigram while also improving the test perplexity and recognition accuracy.
Scalable backoff language models
86 Citations2002Kristie Seymore, Roni Rosenfeld
This project investigates the degradation of a trigram backoff model's perplexity and word error rates as bigram and trigram cutoffs are increased and shows that excluding trigrams and bigrams based on a weighted version of this difference method results in better perplexityand word error rate performance than excluding trigram and bigram based on counts alone.
Comparison of part-of-speech and automatically derived category-based language models for speech recognition
67 Citations2002Thomas Niesler, E. W. D. Whittaker +1 more
This paper compares various category-based language models when used in conjunction with a word-based trigram by means of linear interpolation to find the largest improvement with a model using automatically determined categories.
Statistical language modeling using a variable context length
48 Citations2002Reinhard Kneser
A measure for the quality of variable-length models is developed and a pruning algorithm for the creation of such models is presented, based on this measure, to address the question how the use of a special backing-off distribution can improve the language models.
Language model size reduction by pruning and clustering
48 Citations2000Joshua Goodman, Jianfeng Gao
Novel clustering techniques can be combined with Stolcke pruning to produce the smallest models at a given perplexity, which creates clustered models that are often larger than the unclustered models, but which can be pruned to model that are smaller than unclUSTered models with the same perplexity.
Language modeling with sentence-level mixtures
47 Citations1994Rukmini Iyer, Mari Ostendorf +1 more
Improvements on the pronunciation prefix tree search organization
47 Citations2002F. Alleva, Xuedong Huang +1 more
This work addresses efficiency issues associated with a search organization based on pronunciation prefix trees (PPTs) by presenting a mechanism that eliminates redundant computations in non-reentrant trees, a comparison of two methods for distributing language model probabilities in PPTs, and results on two look ahead pruning strategies.
Speech recognition and the frequency of recently used words
46 Citations1988Roland Kühn
A modification of the Markov approach, which assigns higher probabilities to recently used words, is proposed and tested against a pure Markov model.
Efficient sampling and feature selection in whole sentence maximum entropy language models
32 Citations1999S.F. Chen, Roni Rosenfeld
This paper develops a non-conditional maximum entropy language model which directly models the probability of an entire sentence or utterance, and presents a novel procedure for feature selection, which exploits discrepancies between the existing model and the training corpus.
Objective methods for evaluating synthetic intonation
23 Citations1999Robert A. Clark, Kurt E Dusterhoff
Microbiology SpectrumCombining Statistical and Syntactic Methods in Recognizing Handwritten Sentences
20 Citations1992Rohini K. Srihari, Charlotte M. Baltus
Two statistical methods of applying syntactic constraints to the output of a handwritten word recognizer on input consisting of sentences/phrases based on syntactic categories associated with words are discussed.
Lattice based language models
15 Citations1997Pierre Dupont, Roni Rosenfeld
Lattice based language models, a new language modeling paradigm, are introduced and two primary features to measure the usefulness of each node are proposed: the training-set history count and the smoothed entropy of its prediction.
Research Showcase @ Carnegie Mellon University (Carnegie Mellon University)Exponential Language Models, Logistic Regression, and Semantic Coherence
13 Citations2018Can Cai, Roni Rosenfeld +1 more
Combination of words and word categories in varigram histories
8 Citations1999Reinhard Blasig
A new kind of language model: category/word varigrams that permits a tight integration of word-based and category-based modeling of word sequences and yields a perplexity reduction as compared to a word varigram of the same size.
Integrating detailed information into a language model
6 Citations2002Ruiqiang Zhang, Ezra Black +2 more
A maximum entropy-based approach to language modeling in which both words together with syntactic and semantic tags in the long history are used as a basis for complex linguistic questions to rescore the N-best hypotheses output of the ATRSPREC speech recognition system.
