Large and Diverse Language Models for Statistical Machine Translation
Edinburgh Research Explorer (University of Edinburgh)Published 19 November 2009Open access
Holger Schwenk, Philipp Koehn
Citations38
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
Methods to combine large language models trained from diverse text sources and applies them to a state-ofart French–English and Arabic–English machine translation system are presented.
Abstract
This paper presents methods to combine large language models trained from diverse text sources and applies them to a stateof-art French–English and Arabic–English machine translation system. We show gains of over 2 BLEU points over a strong baseline by using continuous space language models in re-ranking. 1
Keywords
Computer Science
Moses
4,868 Citations2007Philipp Koehn, Richard Zens +12 more
An open-source toolkit for statistical machine translation whose novel contributions are support for linguistically motivated factors, confusion network decoding, and efficient data formats for translation models and language models.
Europarl: A Parallel Corpus for Statistical Machine Translation
3,106 Citations2005Philipp Koehn
A corpus of parallel text in 11 languages from the proceedings of the European Parliament is collected and its acquisition and application as training data for statistical machine translation (SMT) is focused on.
Applied Physics Letters10.1162/153244303322533223
1,771 Citations2000
This work proposes to fight the curse of dimensionality by learning a distributed representation for words which allows each training sentence to inform the model about an exponential number of semantically neighboring sentences.
Large Language Models in Machine Translation
550 Citations2007Thorsten Brants, Ashok C. Popat +3 more
Systems, methods, and computer program products for machine translation are provided for backoff score determination as a function of a backoff factor and a relative frequency of a corresponding backoff n-gram in the corpus.
Computer Speech & LanguageContinuous space language models
503 Citations2006Holger Schwenk
Highly efficient learning algorithms are described that enable the use of training corpora of several hundred million words and it is shown that this approach can be incorporated into a large vocabulary continuous speech recognizer using a lattice rescoring framework at a very low additional processing time.
(Meta-) evaluation of machine translation
385 Citations2007Chris Callison-Burch, Cameron Shaw Fordyce +3 more
An extensive human evaluation was carried out not only to rank the different MT systems, but also to perform higher-level analysis of the evaluation process, revealing surprising facts about the most commonly used methodologies.
Manual and automatic evaluation of machine translation between European languages
276 Citations2006Philipp Koehn, Christof Monz
This work evaluated machine translation performance for six European language pairs that participated in a shared task: translating French, German, Spanish texts to English and back.
Randomised Language Modelling for Statistical Machine Translation
73 Citations2007David Talbot, Miles Osborne
It is shown how a BF containing n-grams can enable us to use much larger corpora and higher-order models complementing a conventional n- gram LM within an SMT system.
Smoothed Bloom Filter Language Models: Tera-Scale LMs on the Cheap
60 Citations2007David Talbot, Miles Osborne
This proposal takes advantage of the one-sided error guarantees of the BF and simple inequalities that hold between related n-gram statistics in order to further reduce the BF storage requirements and the error rate of the derived probabilities.
Distributed language modeling for<i>N</i>-best list re-ranking
54 Citations2006Ying Zhang, Almut Silja Hildebrand +1 more
A novel distributed language model based on the client/server paradigm for N-best list re-ranking that allows for using an arbitrarily large corpus in a very efficient way and provides a natural platform for relevance weighting and selection.
RWTH Publications (RWTH Aachen)Efficient Phrase-Table Representation for Machine Translation with Applications to Online MT and Speech Translation
45 Citations2007Richard Zens, Hermann Ney
A novel algorithm is described that effectively solves this combinatorial problem exploiting the prefix tree data structure of the phrase-table and enables the use of significantly larger input word graphs in a more efficient way resulting in improved translation quality.
Large-Scale Distributed Language Modeling
31 Citations2007Ahmad Emami, Kishore Papineni +1 more
A novel distributed language model that has no constraints on the n-gram order and no practical constraints on vocabulary size is presented, which is scalable and allows for an arbitrarily large corpus to be queried for statistical estimates.
Continuous space language models for the iwslt 2006 task
23 Citations2014Holger Schwenk, Marta R. Costa‐jussà +1 more
This work proposes a new statistical language model that is based on a continuous representation of the words in the vocabulary that is promising for tasks where a very limited amount of resources are available, like the BTEC corpus of tourism related questions.
HAL (Le Centre pour la Communication Scientifique Directe)A state-of-the-art Statistical Machine Translation System based on Moses
8 Citations2007Daniel Déchelotte, Holger Schwenk +3 more
