Leveraging multiple languages to improve statistical MT word alignments
Published 1 January 2005
Karim Filali, Jeffrey A. Bilmes
Citations12
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
A new multilingual statistical MT word alignment model based on a simple extension of the IBM and HMM models and a two-step alignment procedure that shows a 7% relative improvement over a state of the art alignment model.
Abstract
We present a new multilingual statistical MT word alignment model based on a simple extension of the IBM and HMM models and a two-step alignment procedure. Preliminary results on a small hand-aligned subset of the Europarl corpus show a 7% relative improvement over a state of the art alignment model
Keywords
Computer Science
The mathematics of statistical machine translation: parameter estimation
4,125 Citations1993Peter F. Brown, Vincent J. Della Pietra +2 more
It is reasonable to argue that word-by-word alignments are inherent in any sufficiently large bilingual corpus, given a set of pairs of sentences that are translations of one another.
Computational LinguisticsA Systematic Comparison of Various Statistical Alignment Models
3,928 Citations2003Franz Josef Och, Hermann Ney
An important result is that refined alignment models with a first-order dependence and a fertility model yield significantly better results than simple heuristic models.
A statistical approach to machine translation
1,699 Citations1990Peter F. Brown, John Cocke +6 more
The application of the statistical approach to translation from French to English and preliminary results are described and the results are given.
Improved statistical alignment models
1,015 Citations2000Franz Josef Och, Hermann Ney
It is shown that models with a first-order dependence and a fertility model lead to significantly better results than the simple models IBM-1 or IBM-2, which are not able to go beyond zero-order dependencies.
HMM-based word alignment in statistical translation
829 Citations1996Stephan Vogel, Hermann Ney +1 more
A new model for word alignment in statistical translation using a first-order Hidden Markov model for the word alignment problem as they are used successfully in speech recognition for the time alignment problem.
Computational LinguisticsThe mathematics of statistical machine translation
299 Citations1993F BrownPeter, PietraVincent J. Della +2 more
A series of five statistical models of the translation process are described and algorithms for estimating the parameters of these models given a set of pairs of sentences that are translations are given.
A statistical approach to sense disambiguation in machine translation
105 Citations1991Peter F. Brown, Stephen A. Della Pietra +2 more
A statistical technique for assigning senses to words is described and an instance of a word is assigned a sense by asking a question about the context in which the word appears.
Statistical multi-source translation
89 Citations2001Franz Josef Och, Hermann Ney
In various tests, it is shown that these methods can significantly improve translation quality and compare the quality of statistical machine translation systems for many European languages in the same domain.
Extensions to HMM-based statistical word alignment models
88 Citations2002Kristina Toutanova, H. Tolga Ilhan +1 more
Improved HMM-based word level alignment models for statistical machine translation and a method for using part of speech tag information to improve alignment accuracy, and an approach to modeling fertility and correspondence to the empty word in an HMM alignment model.
Shared task
70 Citations2005Philipp Koehn, Christof Monz
The goals, the task definition and resources, as well as results and some analysis are described, which describe a shared task on building statistical machine translation systems for four European language pairs.
Word alignment for languages with scarce resources
69 Citations2005Joel Martin, Rada Mihalcea +1 more
The task definition, resources, participating systems, and comparative results for the shared task on word alignment, which was organized as part of the ACL 2005 Workshop on Building and Using Parallel Texts, are presented.
Text-Translation Alignment: Three Languages Are Better Than Two.
43 Citations1999Michel Simard
Experiments on a trilingual corpus demonstrate that this bilingual texttranslation alignment method can be adapted to deal with more than two versions of a text and yields better bilingual alignments than can be obtained with bilingual textalignment methods.
Improved HMM alignment models for languages with scarce resources
24 Citations2005Adam Lopez, Philip Resnik
Improvements to statistical word alignment based on the Hidden Markov Model incorporate syntactic knowledge and show that alignment performance exceeds that of a state-of-the art system based on more complex models.
Bilingual word spectral clustering for statistical machine translation
10 Citations2005Bing Zhao, Eric P. Xing +1 more
A variant of a spectral clustering algorithm is proposed for bilingual word clustering that generates the two sets of clusters for both languages efficiently with high semantic correlation within monolingual clusters, and high translation quality across the clusters between two languages.
