Is Machine Translation Getting Better over Time?
Generate an AI Snapshot to get a quick, structured summary of this paper.
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
A large-scale crowd-sourcing experiment is carried out to estimate the degree to which state-of-theart performance in machine translation has increased over the past five years, with Czech-to-English translation standing out as the language pair achieving most substantial gains.
Abstract
Recent human evaluation of machine translation has focused on relative pref-erence judgments of translation quality, making it difficult to track longitudinal im-provements over time. We carry out a large-scale crowd-sourcing experiment to estimate the degree to which state-of-the-art performance in machine translation has increased over the past five years. To fa-cilitate longitudinal evaluation, we move away from relative preference judgments and instead ask human judges to provide direct estimates of the quality of individ-ual translations in isolation from alternate outputs. For seven European language pairs, our evaluation estimates an aver-age 10-point improvement to state-of-the-art machine translation between 2007 and 2012, with Czech-to-English translation standing out as the language pair achiev-ing most substantial gains. Our method of human evaluation offers an economi-cally feasible and robust means of per-forming ongoing longitudinal evaluation of machine translation. 1
