DCU-Lingo24 Participation in WMT 2014 Hindi-English Translation task
Published 1 January 2014Open access
Xiaofeng Wu, Rejwanul Haque, Tsuyoshi Okita, Piyush Arora, Andy Way, Qun Liu
Citations2
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
The DCU-Lingo24 submission to WMT 2014 for the HindiEnglish translation task is described and miscellaneous methods in the system, including: Context-Informed PB-SMT, OOV Word Conversion, MultiAlignment Combination, Operation Sequence Model, Stemming Align and Normal Phrase Extraction, and Language Model Interpolation are exploited.
Abstract
This paper describes the DCU-Lingo24 submission to WMT 2014 for the Hindi-English translation task. We exploit miscellaneous methods in our system,
Keywords
Computer Science
Moses
4,868 Citations2007Philipp Koehn, Richard Zens +12 more
An open-source toolkit for statistical machine translation whose novel contributions are support for linguistically motivated factors, confusion network decoding, and efficient data formats for translation models and language models.
Computational LinguisticsA Systematic Comparison of Various Statistical Alignment Models
3,928 Citations2003Franz Josef Och, Hermann Ney
An important result is that refined alignment models with a first-order dependence and a fertility model yield significantly better results than simple heuristic models.
Statistical phrase-based translation
3,268 Citations2003Philipp Koehn, Franz Josef Och +1 more
The empirical results suggest that the highest levels of performance can be obtained through relatively simple means: heuristic learning of phrase translations from word-based alignments and lexical weighting of phrase translation.
Minimum error rate training in statistical machine translation
2,766 Citations2003Franz Josef Och
It is shown that significantly better results can often be obtained if the final evaluation criterion is taken directly into account as part of the training procedure.
arXiv (Cornell University)Exploiting Similarities among Languages for Machine Translation
1,435 Citations2013Tomáš Mikolov, Quoc V. Le +1 more
This method can translate missing word and phrase entries by learning language structures based on large monolingual data and mapping between languages from small bilingual data and uses distributed representation of words and learns a linear mapping between vector spaces of languages.
A Neural Probabilistic Language Model
1,156 Citations2000Yoshua Bengio, Réjean Ducharme +1 more
FigshareZero-Shot Learning with Semantic Output Codes
829 Citations2018Mark Palatucci, Dean Pomerleau +2 more
A semantic output code classifier which utilizes a knowledge base of semantic properties of Y to extrapolate to novel classes and can often predict words that people are thinking about from functional magnetic resonance images of their neural activity, even without training examples for those words.
Applying Conditional Random Fields to Japanese Morphological Analysis
723 Citations2004Taku Kudo, Kaoru Yamamoto +1 more
This paper shows how CRFs can be applied to situations where word boundary ambiguity exists, and confirms that CRFs offer a solution to the long-standing problems in corpus-based or statistical Japanese morphological analysis.
ArXiv.orgAn Empirical Study of Smoothing Techniques for Language Modeling
633 Citations1996Stanley F. Chen, Joshua Goodman
Alignment by agreement
433 Citations2006Percy Liang, Ben Taskar +1 more
An unsupervised approach to symmetric word alignment in which two simple asymmetric models are trained jointly to maximize a combination of data likelihood and agreement between the models.
Edinburgh Research Explorer (University of Edinburgh)Edinburgh System Description for the 2005 IWSLT Speech Translation Evaluation
366 Citations2005Philipp Koehn, Amittai Axelrod +4 more
This work adapted the statistical machine translation system that performed successfully in previous DARPA competitions on open domain text translations to work on limited domain speech data in the IWSLT 2005 speech translation task.
A simple and effective hierarchical phrase reordering model
297 Citations2008Michel Galley, Christopher D. Manning
A novel hierarchical phrase reordering model aimed at improving non-local reorderings, which seamlessly integrates with a standard phrase-based system with little loss of computational efficiency is presented.
Dependency Annotation Scheme for Indian Languages
118 Citations2008Rafiya Begum, Samar Husain +4 more
The motivation for following thePaninian framework as the annotation scheme is provided and it is argued that the Paninian framework is better suited to model the various linguistic phenomena manifest in Indian languages.
A Joint Sequence Translation Model with Integrated Reordering
108 Citations2011Nadir Durrani, Helmut Schmid +1 more
A novel machine translation model which models translation by a linear sequence of operations which includes not only translation but also reordering operations, and a joint sequence model for the translation and reordering probabilities which is more flexible than standard phrase-based MT.
Computers and the HumanitiesUnsupervised morphological parsing of Bengali
58 Citations2007Sajib Dasgupta, Vincent Ng
This paper introduces a simple, yet highly effective algorithm for unsupervised morphological learning for Bengali, an Indo–Aryan language that is highly inflectional in nature.
Combining Multiple Alignments to Improve Machine Translation
15 Citations2012Zhaopeng Tu, Yang Liu +4 more
Experiments show that this approach not only improves the alignment quality, but also significantly improves translation performance by up to 1.96 BLEU points over single best alignments, and 1.28 points over merging rules extracted from multiple alignments individually.
Data Issues in English-to-Hindi Machine Translation
13 Citations2010Ondřej Bojar, Pavel Straňák +1 more
This paper discusses several available parallel data sources and provides cross-evaluation results on their combinations using two freely available statistical MT systems and presents a new tool for viewing aligned corpora, which makes it easier to detect difficult parts in the data even for a developer not speaking the target language.
