Findings of the 2014 Workshop on Statistical Machine Translation
Published 1 January 2014Open access
Ondřej Bojar, Christian Buck, Christian Federmann, Barry Haddow, Philipp Koehn, Johannes Leveling
Citations527
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
The results of the WMT14 shared tasks, which included a standard news translation task, a separate medical translationtask, a task for run-time estimation of machine translation quality, and a metrics task, are presented.
Abstract
Ondřej Bojar, Christian Buck, Christian Federmann, Barry Haddow, Philipp Koehn, Johannes Leveling, Christof Monz, Pavel Pecina, Matt Post, Herve Saint-Amand, Radu Soricut, Lucia Specia, Aleš Tamchyna. Proceedings of the Ninth Workshop on Statistical Machine Translation. 2014.
Keywords
Computer ScienceBiochemistry, Genetics and Molecular Biology
BiometricsThe Measurement of Observer Agreement for Categorical Data
78,934 Citations1977J. Richard Landis, Gary G. Koch
A general statistical methodology for the analysis of multivariate categorical data arising from observer reliability studies is presented and tests for interobserver bias are presented in terms of first-order marginal homogeneity and measures of interob server agreement are developed as generalized kappa-type statistics.
Machine LearningExtremely randomized trees
8,669 Citations2006Pierre Geurts, Damien Ernst +1 more
A new tree-based ensemble method for supervised classification and regression problems that consists of randomizing strongly both attribute and cut-point choice while splitting a tree node and builds totally randomized trees whose structures are independent of the output values of the learning sample.
ACM Transactions on Information SystemsCumulated gain-based evaluation of IR techniques
4,634 Citations2002Kalervo Järvelin, Jaana Kekäläinen
This article proposes several novel measures that compute the cumulative gain the user obtains by examining the retrieval result up to a given ranked position, and test results indicate that the proposed measures credit IR methods for their ability to retrieve highly relevant documents and allow testing of statistical significance of effectiveness differences.
arXiv (Cornell University)Evaluation: from precision, recall and F-measure to ROC, informedness, markedness and correlation
4,425 Citations2020David Powers
E elegant connections between the concepts of Informedness, Markedness, Correlation and Significance as well as their intuitive relationships with Recall and Precision are demonstrated.
A Study of Translation Edit Rate with Targeted Human Annotation
2,402 Citations2006Matthew Snover, Bonnie J. Dorr +3 more
A new, intuitive measure for evaluating machine translation output that avoids the knowledge intensiveness of more meaning-based approaches, and the labor-intensiveness of human judgments is defined.
Discriminative training methods for hidden Markov models
1,889 Citations2002Michael Collins
Experimental results on part-of-speech tagging and base noun phrase chunking are given, in both cases showing improvements over results for a maximum-entropy tagger.
Nucleic Acids ResearchDrugBank 3.0: a comprehensive resource for 'Omics' research on drugs
1,857 Citations2010Craig Knox, Vivian Law +12 more
DrugBank 3.0 represents the result of 2 years of manual annotation work aimed at making the database much more useful for a wide range of ‘omics’ applications, particularly with regard to drug target, drug description and drug action data.
BioinformaticsGENIA corpus—a semantically annotated corpus for bio-textmining
1,243 Citations2003JD Kim, Tomoko Ohta +2 more
Retrieval evaluation with incomplete information
733 Citations2004Chris Buckley, Ellen M. Voorhees
It is shown that current evaluation measures are not robust to substantially incomplete relevance judgments, and a new measure is introduced that is both highly correlated with existing measures when complete judgments are available and more robust to incomplete judgment sets.
The MIT Press eBooksTrueSkill™: A Bayesian Skill Rating System
625 Citations2007Ralf Herbrich, Tom Minka +1 more
A new Bayesian skill rating system which can be viewed as a generalisation of the Elo system used in Chess, which tracks the uncertainty about player skills, explicitly models draws, can deal with any number of competing entities and can infer individual skills from team results.
Journal of Machine Learning TechnologiesJournal of Machine Learning Technologies
525 Citations2024
The Journal of Machine Learning Technologies, which contains a series of timely, in-depth written articles by leaders in the field, covering a wide range of the integration of multidimensional challenges of research including integration issues of Machine learning Technologie.
(Meta-) evaluation of machine translation
385 Citations2007Chris Callison-Burch, Cameron Shaw Fordyce +3 more
An extensive human evaluation was carried out not only to rank the different MT systems, but also to perform higher-level analysis of the evaluation process, revealing surprising facts about the most commonly used methodologies.
Findings of the 2009 workshop on statistical machine translation
351 Citations2009Chris Callison-Burch, Philipp Koehn +2 more
A large-scale manual evaluation of 103 machine translation systems submitted by 34 teams was conducted, which used the ranking of these systems to measure how strongly automatic metrics correlate with human judgments of translation quality for 12 evaluation metrics.
Findings of the 2013 Workshop on Statistical Machine Translation
298 Citations2013Chris Callison-Burch, Philipp Koehn +6 more
Manual and automatic evaluation of machine translation between European languages
276 Citations2006Philipp Koehn, Christof Monz
This work evaluated machine translation performance for six European language pairs that participated in a shared task: translating French, German, Spanish texts to English and back.
Further meta-evaluation of machine translation
269 Citations2008Chris Callison-Burch, Cameron Shaw Fordyce +3 more
This paper analyzes the translation quality of machine translation systems for 10 language pairs translating between Czech, English, French, German, Hungarian, and Spanish and uses the human judgments of the systems to analyze automatic evaluation metrics for translation quality.
Institutional Research Information System (Università degli Studi di Trento)Findings of the 2013 Workshop on Statistical Machine Translation
243 Citations2013Ondřej Bojar, Christian Buck +8 more
The results of the WMT13 shared tasks, which included a translation task, a task for run-time estimation of machine translation quality, and an unofficial metrics task are presented.
Accelerated DP based search for statistical translation
238 Citations1997Christoph Tillmann, S. Vogel +3 more
A fast search algorithm for statistical translation based on dynamic programming (DP) based on the assumption that the word alignment is monotone with respect to the word order in both languages is described and results are presented.
Computational biologyThe Foundational Model of Anatomy Ontology
232 Citations2007Cornelius Rosse, José L. V. Mejino
The Foundational Model of Anatomy (FMA) ontology is being developed to fill the need for a generalizable anatomy ontology, which can be used and adapted by any computer-based application that requires anatomical information.
Findings of the 2010 Joint Workshop on Statistical Machine Translation and Metrics for Machine Translation
185 Citations2010Chris Callison-Burch, Philipp Koehn +4 more
A large-scale manual evaluation of 104 machine translation systems and 41 system combination entries was conducted, which used the ranking of these systems to measure how strongly automatic metrics correlate with human judgments of translation quality for 26 metrics.
White Rose Research Online (University of Leeds, The University of Sheffield, University of York)QuEst - A translation quality estimation framework
180 Citations2013Lucia Specia, Kashif Shah +2 more
quest, an open source framework for machine translation quality estimation that allows the extraction of several quality indicators from source segments, their translations, external resources, as well as language tools, and provides machine learning algorithms to build quality estimation models.
Meeting of the Association for Computational LinguisticsTranslationese and Its Dialects
137 Citations2011Moshe Koppel, Noam Ordan
It is found that even for widely unrelated source languages and multiple genres, differences between translated texts and non-translated texts are sufficient for a learned classifier to accurately determine if a given text is translated or original.
BMC BioinformaticsConstruction of an annotated corpus to support biomedical information extraction
122 Citations2009Paul M. Thompson, Syed Amir Iqbal +2 more
A new scheme for annotating sentence-bound gene regulation events, centred on both verbs and nominalised verbs is defined, and the GREC is created, consisting of 240 MEDLINE abstracts, in which events relating to gene regulation and expression have been annotated by biologists.
Conference of the European Chapter of the Association for Computational LinguisticsCDER: Efficient MT Evaluation Using Block Movements.
111 Citations2006Gregor Leusch, Nicola Ueffing +1 more
Assessment of the bacteriology and antibiotic susceptibility of breast implant-related infections at two tertiary care hospitals in the Texas Medical Center found that choosing an antibiotic with anti-methicillin-resistant S. aureus activity is justified for empiric treatment of Breast implant infections, until culture and sensitivity data, if obtained, become available.
Results of the WMT14 Metrics Shared Task
109 Citations2014Matouš Macháċek, Ondřej Bojar
This paper presents the results of the WMT14 Metrics Shared Task, which asked participants of this task to score the outputs of the MT systems involved in W MT14 Shared Translation Task to evaluate system level correlation and segment level correlation.
Modelling Annotator Bias with Multi-task Gaussian Processes: An Application to Machine Translation Quality Estimation
106 Citations2013Trevor Cohn, Lucia Specia
Novel techniques for learning from the outputs of multiple annotators while accounting for annotator specific behaviour are presented, which use multi-task Gaussian Processes to learn jointly a series of annotator and metadata specific models.
Efficient Elicitation of Annotations for Human Evaluation of Machine Translation
57 Citations2014Keisuke Sakaguchi, Matt Post +1 more
The experimental results show that TrueSkill outperforms other recently proposed models on accuracy, and also can significantly reduce the number of pairwise annotations that need to be collected by sampling non-uniformly from the space of system competitions.
A Grain of Salt for the WMT Manual Evaluation
50 Citations2011Ondřej Bojar, Miloš Ercegovċević +2 more
The aim of this paper is to investigate and explain some interesting idiosyncrasies in the reported results, which only become apparent when performing a more thorough analysis of the collected annotations.
FBK-UPV-UEdin participation in the WMT14 Quality Estimation shared-task
45 Citations2014José Guilherme Camargo de Souza, Jesús González-Rubio +3 more
The joint submission of Fondazione Bruno Kessler, Universitat Politde Val` encia and University of Edinburgh to the Quality Estimation tasks of the Workshop on Statistical Machine Translation 2014 is described.
Edinburgh’s Phrase-based Machine Translation Systems for WMT-14
44 Citations2014Nadir Durrani, Barry Haddow +2 more
UEDIN’s phrase-based submissions to the translation and medical translation shared tasks of the 2014 Workshop on Statistical Machine Translation (WMT) are described.
An Investigation on the Effectiveness of Features for Translation Quality Estimation
41 Citations2013Kashif Shah, Trevor Conn +1 more
Phrasal: A Toolkit for New Directions in Statistical Machine Translation
40 Citations2014Spence Green, Daniel Cer +1 more
A new version of Phrasal, an open-source toolkit for statistical phrasebased machine translation, is presented, which includes features that support emerging research trends such as tuning with large feature sets, and web-based interactive machine translation.
The IIT Bombay Hindi-English Translation System at WMT 2014
34 Citations2014Piyush Dungarwal, Rajen Chatterjee +4 more
It is shown that the use of number, case and Tree Adjoining Grammar information as factors helps to improve English-Hindi translation, primarily by generating morphological inflections correctly.
Information Design JournalIntegrating content and style in documents
32 Citations2000Nadjet Bouayad‐Agha, Donia R. Scott +1 more
Learning syntactic structure
25 Citations2007Yoav Seginer
This dissertation aims to provide a history of web exceptionalism from 1989 to 2002, a period chosen in order to explore its roots as well as specific cases up to and including the year in which descriptions of “Web 2.0” began to circulate.
LIG System for Word Level QE task at WMT14
24 Citations2014Ngoc Quang Luong, Laurent Besacier +1 more
The Word-level QE system for WMT 2014 shared task on Spanish-English pair is described, optimized by several ways: tuning the classification threshold, combining with WMT 2013 data, and refining using Feature Selection strategy on the development set, before dealing with the test set for submission.
Manawi: Using Multi-Word Expressions and Named Entities to Improve Machine Translation
23 Citations2014Liling Tan, Santanu Pal
The Manawi system showed the potential of improving translation quality by incorporating multiple NLP tools within the MT pipeline by introducing a novel filter method based on sentence-alignment features.
Lecture notes in computer scienceAnalyzing Parallelism and Domain Similarities in the MAREC Patent Corpus
22 Citations2012Katharina Wäschle, Stefan Riezler
A twofold approach for extracting parallel data from all patent document sections from a large multilingual patent corpus and a descriptive analysis of its subdomains to enable its use in domain-oriented translation, e.g. when applying multi-task learning.
Referential Translation Machines for Predicting Translation Quality
19 Citations2014Ergun Biçici, Andy Way
Ref referential translation machines remove the need to access any SMT system specific information or prior knowledge of the training data or models used when generating the translations and achieve the top performance in WMT13 quality estimation task (QET13).
Machine Translation of Medical Texts in the Khresmoi Project
19 Citations2014Ondřej Dušek, Jan Hajič +7 more
The participation of the Charles University team in the WMT 2014 Medical Translation Task is presented, with a primary goal to set up a baseline for both its subtasks and for all translation directions.
EU-BRIDGE MT: Combined Machine Translation
18 Citations2014Markus Freitag, Stephan Peitz +11 more
Three research institutes involved in the EU-BRIDGE project combined their individual machine translation systems and participated with a joint setup in the shared translation task of the evaluation campaign at the ACL 2014 Eighth Workshop on Statistical Machine Translation.
Edinburgh Research Explorer (University of Edinburgh)Putting Human Assessments of Machine Translation Systems in Order
18 Citations2012Adam Lopez
This work extends their analysis to all of the ranking tasks from 2010 and 2011, and shows that the ranking is naturally cast as an instance of finding the minimum feedback arc set in a tournament, a well-known NP-complete problem.
LIG System for WMT13 QE Task: Investigating the Usefulness of Features in Word Confidence Estimation for MT
18 Citations2013Ngoc Quang Luong, Benjamin Lecouteux +1 more
This paper presents the LIG’s systems submitted for Task 2 of WMT13 Quality Estimation campaign, a word confidence estimation task where each participant was asked to label each word in a translated text as a binary or multi-class category.
Edinburgh Research Explorer (University of Edinburgh)Simulating Human Judgment in Machine Translation Evaluation Campaigns
18 Citations2012Philipp Koehn
A Monte Carlo model is presented to simulate human judgments in machine translation evaluation campaigns, such as WMT or IWSLT, to compare different ranking methods and to give guidance on the number of judgments that need to be collected to obtain sufficiently significant distinctions between systems.
Parallel FDA5 for Fast Deployment of Accurate Statistical Machine Translation Systems
16 Citations2014Ergun Biçici, Qun Liu +1 more
This work builds Parallel FDA5 Moses SMT systems for all language pairs in the WMT14 translation task and obtains SMT performance close to the top Moses systems with an average of 3.49 BLEU points difference using significantly less resources for training and development.
The CMU Machine Translation Systems at WMT 2014
13 Citations2014Austin Matthews, Waleed Ammar +8 more
Inventions include: a label coarsening scheme for syntactic tree-to-tree translation, a host of new discriminative features, several modules to create “synthetic translation options” that can generalize beyond what is directly observed in the training data, and a method of combining the output of multiple word aligners to uncover extra phrase pairs and grammar rules.
Meeting of the Association for Computational LinguisticsModels of Translation Competitions
13 Citations2013Mark Hopkins, Jonathan May
This work provides the first framework that allows an empirical comparison of different analyses of competition results, and uses this framework to compare several analytical models on data from the Workshop on Machine Translation (WMT).
Quality estimation for translation selection.
12 Citations2014Kashif Shah, Lucia Specia
Machine Translation and Monolingual Postediting: The AFRL WMT-14 System
11 Citations2014Lane Schwartz, Timothy R. Anderson +2 more
The AFRL statistical MT system and the improvements that were developed during the WMT14 evaluation campaign are described and the efforts to make use of monolingual English speakers to correct the output of machine translation are described.
LIMSI Submission for WMT'14 QE Task
10 Citations2014Guillaume Wisniewski, Nicolas Pécheux +2 more
LIMSI participation to the WMT’14 Shared Task on Quality Estimation; the system relies on a random forest classifier, an ensemble method that has been shown to be very competitive for this kind of task, when only a few dense and continuous features are used.
Yandex School of Data Analysis Russian-English Machine Translation System for WMT14
9 Citations2014Alexey Borisov, Irina Galinskaya
This paper describes the Yandex School of Data Analysis Russian-English system and proposes a {simple yet practical} algorithm to transform Russian sentence into a more easily translatable form before decoding.
The DCU-ICTCAS MT system at WMT 2014 on German-English Translation Task
9 Citations2014Liangyou Li, Xiaofeng Wu +4 more
This paper describes the DCU submission to WMT 2014 on German-English translation task, which uses phrasebased translation model with several popular techniques, including Lexicalized Reordering Model, Operation Sequence Model and Language Model interpolation.
Exploiting Qualitative Information from Automatic Word Alignment for Cross-lingual NLP Tasks
9 Citations2013José G. C. de Souza, Miquel Esplà-Gomis +2 more
This paper contributes with a novel method that significantly outperforms the state of the art, and is portable, with limited loss in performance, to language pairs where training data are not available.
SHEF-Lite 2.0: Sparse Multi-task Gaussian Processes for Translation Quality Estimation
9 Citations2014Daniel Beck, Kashif Shah +1 more
These submissions use the framework of Multi-task Gaussian Processes, where they combine multiple datasets in a multi-task setting to speed up training and prediction by providing sensible sparse approximations.
Combining Domain Adaptation Approaches for Medical Text Translation
9 Citations2014Longyue Wang, Yi Lü +4 more
A number of simple and effective techniques to adapt statistical machine translation systems in the medical domain and these systems achieve the best BLEU scores for Czech-English, EnglishGerman, French-English language pairs and the second best Blemish scores for reminding pairs are explored.
The RWTH Aachen German-English Machine Translation System for WMT 2014
8 Citations2014Stephan Peitz, Joern Wuebker +2 more
Anaphora Models and Reordering for Phrase-Based SMT
8 Citations2014Christian Hardmeier, Sara Stymne +3 more
The Uppsala University systems for WMT14 are described and the integration of a model for translating pronominal anaphora and a syntactic dependency projection model for English‐French and tunable POS distortion models for English- German are investigated.
LIMSI $@$ WMT’14 Medical Translation Task
7 Citations2014Nicolas Pécheux, Li Gong +8 more
LIMSI’s submission to the first medical translation task at WMT’14 is described and results for EnglishFrench on the subtask of sentence translation from summaries of medical articles are reported.
Domain Adaptation for Medical Text Translation using Web Resources
7 Citations2014Yi Lü, Longyue Wang +3 more
This paper describes adapting statistical machine translation systems to medical domain using in-domain and general-domain data as well as webcrawled in- domain resources and proposes an alternative filtering approach to clean the crawled data and to further optimize the domain-specific SMT system.
CUNI in WMT14: Chimera Still Awaits Bellerophon
7 Citations2014Aleš Tamchyna, Martin Popel +2 more
The English!Czech and English!Hindi submissions for this year’s WMT translation task are presented and reverse self-training to acquire more parallel data and with modeling target-side morphology are experimented with.
Experiments in Medical Translation Shared Task at WMT 2014
6 Citations2014Jian Zhang
Dublin City University’s (DCU) submission to the WMT 2014 Medical Summary task is described and results on the test data set in the French to English translation direction are reported.
CimS – The CIS and IMS joint submission to WMT 2014 translating from English into German
6 Citations2014Fabienne Cap, Marion Weller +2 more
The morphologyaware translation systems handle word formation issues on different levels of morpho-syntactic modeling, including complex nominal and verbal morphology, productive compounding and flexible word ordering.
The Karlsruhe Institute of Technology Translation Systems for the WMT 2014
6 Citations2014Teresa Herrmann, Mohammed Mediani +6 more
A discriminative word lexicon using source context information proved beneficial for all translation directions, and was applied to make use of noisy web-crawled data.
Abu-MaTran at WMT 2014 Translation Task: Two-step Data Selection and RBMT-Style Synthetic Rules
5 Citations2014Raphaël Rubino, Antonio Toral +6 more
The UA-Prompsit hybrid machine translation system for the 2014 Workshop on Statistical Machine Translation
5 Citations2014Víctor M. Sánchez-Cartagena, Juan Antonio Pérez-Ortiz +1 more
A phrase-based statistical machine translation system whose phrase table is enriched with information obtained from dictionaries and shallowtransfer rules like those used in rule-based machine translation.
Target-Centric Features for Translation Quality Estimation
5 Citations2014Chris Hokamp, Iacer Calixto +2 more
The DCU-MIXED and DCUSVR submissions to the WMT-14 Quality Estimation task 1.1 are described, predicting sentencelevel perceived post-editing effort and feature design focuses on target-side features as it is hypothesised that the source side has little effect on the quality of human translations.
Exploring Consensus in Machine Translation for Quality Estimation
4 Citations2014Carolina Scarton, Lucia Specia
The use of consensus among Machine Translation (MT) systems for the WMT14 Quality Estimation shared task is presented by comparing the MT system output against several alternative machine translations using standard evaluation metrics.
Large-scale Exact Decoding: The IMS-TTT submission to WMT14
3 Citations2014Daniel Quernheim, Fabienne Cap
This work presents the IMS-TTT submission to WMT14, an experimental statistical treeto-tree machine translation system based on the multi-bottom up tree transducer including rule extraction, tuning and decoding, and the obtained translations are competitive.
Postech's System Description for Medical Text Translation Task
3 Citations2014Jianri Li, Se‐Jong Kim +2 more
This short paper presents a system description for intrinsic evaluation of the WMT 14’s medical text translation task using phrase-based statistical machine translation system and query translation system between German-English language pairs.
DCU-Lingo24 Participation in WMT 2014 Hindi-English Translation task
2 Citations2014Xiaofeng Wu, Rejwanul Haque +4 more
The DCU-Lingo24 submission to WMT 2014 for the HindiEnglish translation task is described and miscellaneous methods in the system, including: Context-Informed PB-SMT, OOV Word Conversion, MultiAlignment Combination, Operation Sequence Model, Stemming Align and Normal Phrase Extraction, and Language Model Interpolation are exploited.
The KIT-LIMSI Translation System for WMT 2014
1 Citations2014Quoc Khanh, Teresa Herrmann +4 more
Experimental results show that SOUL translation models use in the KIT phrase-based system can yield significant improvements in terms of BLEU score.
…
