Large, Pruned or Continuous Space Language Models on a GPU for Statistical Machine Translation
HAL (Le Centre pour la Communication Scientifique Directe)Published 1 January 2012
Holger Schwenk, Anthony Rousseau, Mohammed Attik
Citations90
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
An experimental comparison of all the approaches to train and use continuous space language models (CSLM) on a large statistical machine translation task and an efficient implementation of the CSLM using graphical processing units from Nvidia is described.
Abstract
International audience
Keywords
Computer Science
Recurrent neural network based language model
5,349 Citations2010Tomáš Mikolov, Martin Karafiát +3 more
Results indicate that it is possible to obtain around 50% reduction of perplexity by using mixture of several RNN LMs, compared to a state of the art backoff language model.
Moses
4,868 Citations2007Philipp Koehn, Richard Zens +12 more
An open-source toolkit for statistical machine translation whose novel contributions are support for linguistically motivated factors, confusion network decoding, and efficient data formats for translation models and language models.
Computer Speech & LanguageAn empirical study of smoothing techniques for language modeling
2,077 Citations1999Stanley F. Chen, Joshua Goodman
Applied Physics Letters10.1162/153244303322533223
1,771 Citations2000
This work proposes to fight the curse of dimensionality by learning a distributed representation for words which allows each training sentence to inform the model about an exponential number of semantically neighboring sentences.
Extensions of recurrent neural network language model
1,596 Citations2011Tomáš Mikolov, Stefan Kombrink +3 more
Several modifications of the original recurrent neural network language model are presented, showing approaches that lead to more than 15 times speedup for both training and testing phases and possibilities how to reduce the amount of parameters in the model.
A Neural Probabilistic Language Model
1,156 Citations2000Yoshua Bengio, Réjean Ducharme +1 more
KenLM: Faster and Smaller Language Model Queries
1,142 Citations2011Kenneth Heafield
KenLM is a library that implements two data structures for efficient language model queries, reducing both time and memory costs and is integrated into the Moses, cdec, and Joshua translation systems.
A Scalable Hierarchical Distributed Language Model
852 Citations2008Andriy Mnih, Geoffrey E. Hinton
A fast hierarchical language model along with a simple feature-based algorithm for automatic construction of word trees from the data are introduced and it is shown that the resulting models can outperform non-hierarchical neural models as well as the best n-gram models.
International Conference on Artificial Intelligence and StatisticsHierarchical Probabilistic Neural Network Language Model.
840 Citations2005Fréderic Morin, Yoshua Bengio
A point-of-care POC analyzer that can rapidly measure both medications used for treatment and illicit drugs in patient saliva and shows evidence of accuracy in measuring treatment medications and relapse to drug use is shown.
Large Language Models in Machine Translation
550 Citations2007Thorsten Brants, Ashok C. Popat +3 more
Systems, methods, and computer program products for machine translation are provided for backoff score determination as a function of a backoff factor and a relative frequency of a corresponding backoff n-gram in the corpus.
Intelligent Selection of Language Model Training Data
539 Citations2010Robert C. Moore, William D. Lewis
This work addresses the problem of selecting non-domain-specific language model training data to build auxiliary language models for use in tasks such as machine translation by comparing the cross-entropy, according to domain-specific and non- domain-specifc language models, for each sentence of the text source used to produce the latter language model.
Computer Speech & LanguageContinuous space language models
503 Citations2006Holger Schwenk
Highly efficient learning algorithms are described that enable the use of training corpora of several hundred million words and it is shown that this approach can be incorporated into a large vocabulary continuous speech recognizer using a lattice rescoring framework at a very low additional processing time.
Structured Output Layer neural network language model
131 Citations2011Hai-Son Le, Ilya Oparin +3 more
A new neural network language model (NNLM) based on word clustering to structure the output vocabulary: Structured Output Layer NNLM, able to handle vocabularies of arbitrary size, hence dispensing with the design of short-lists that are commonly used in NNLMs.
Continuous space language models for statistical machine translation
127 Citations2006Holger Schwenk, Daniel Dchelotte +1 more
This work proposes to use a new statistical language model that is based on a continuous representation of the words in the vocabulary, which achieves consistent improvements in the BLEU score on the development and test data.
IEEE International Conference on Acoustics Speech and Signal ProcessingConnectionist language modeling for large vocabulary continuous speech recognition
116 Citations2002Holger Schwenk, Jean‐Luc Gauvain
The connectionist language model is being evaluated on the DARPA HUB5 conversational telephone speech recognition task and preliminary results show consistent improvements in both perplexity and word error rate.
Improved neural network based language modelling and adaptation
81 Citations2010Junho Park, Xunying Liu +2 more
A novel NNLM adaptation method using a cascaded network is proposed and consistent WER reductions were obtained on a state-of-the-art Arabic LVCSR task over conventional NNLMs.
Efficient handling of<i>N</i>-gram language models for statistical machine translation
74 Citations2007Marcello Federico, Mauro Cettolo
This paper presents efficient algorithmic and architectural solutions which have been tested within the Moses decoder, an open source toolkit for statistical machine translation, and shows that the proposed implementation seems to scale-up much better to very large language models.
The Prague Bulletin of Mathematical LinguisticsContinuous-Space Language Models for Statistical Machine Translation
71 Citations2010Holger Schwenk
An open-source implementation of the so-called continuous space language model and its application to statistical machine translation and its implementation details and experimental results on a variety of tasks are given.
Smoothed Bloom Filter Language Models: Tera-Scale LMs on the Cheap
60 Citations2007David Talbot, Miles Osborne
This proposal takes advantage of the one-sided error guarantees of the BF and simple inequalities that hold between related n-gram statistics in order to further reduce the BF storage requirements and the error rate of the derived probabilities.
Efficient training of large neural networks for language modeling
31 Citations2005Holger Schwenk
The described approach achieves significant word error reductions with respect to a carefully tuned 4-gram backoff language model in a state of the art conversational speech recognizer for the DARPA rich transcriptions evaluations.
Empirical Methods in Natural Language ProcessingTraining Continuous Space Language Models: Some Practical Issues
30 Citations2010Hai Son Le, Alexandre Allauzen +2 more
This work studies the performance and behavior of two neural statistical language models so as to highlight some important caveats of the classical training algorithms, and introduces a new initialization scheme and new training techniques to greatly reduce the training time and to significantly improve performance.
Efficient Subsampling for Training Complex Language Models
25 Citations2011Puyang Xu, Asela Gunawardana +1 more
It is shown that the binarized model is as powerful as the standard model and allows us to aggressively subsample negative training examples without sacrificing predictive performance.
Large vocabulary SOUL neural network language models
24 Citations2011Hai Son Le, Ilya Oparin +4 more
A new training scheme is proposed for SOUL NNLMs that is based on separate training of the outof-shortlist part of the output layer, which enables using more data at each iteration of a neural network without any considerable slow-down in training and brings additional improvements in speech recognition performance.
Improved models for Mandarin speech-to-text transcription
23 Citations2011Lori Lamel, Jean‐Luc Gauvain +3 more
The improved system reduces the relative character error rate (CER) by about 10% on previous GALE development and evaluation data sets, obtaining a CER of 9.2% on the P4 broadcast news and broadcast conversation evaluation data.
Data selection and smoothing in an open-source system for the 2008 NIST machine translation evaluation
12 Citations2008Holger Schwenk, Yannick Estève
This paper gives a detailed description of a statistical machine translation system developed for the 2008 NIST open MT evaluation based on the open source toolkit Moses with extensions for language model rescoring in a second pass.
Improving LVCSR system combination using neural network language model cross adaptation
11 Citations2011Xiaobing Liu, Mark Gales +1 more
Previous research on multi-level n-gram LM cross adaptation is extended to further include the cross adaptation of neural network LMs in this paper, and significant error rate gains were obtained when combining a range of Chinese LVCSR sub-systems used in the 2010 and 2011 DARPA GALE evaluations.
