deltaBLEU: A Discriminative Metric for Generation Tasks with Intrinsically Diverse Targets
arXiv (Cornell University)Published 23 June 2015Open access
Michel Galley, Chris Brockett, Alessandro Sordoni, Yangfeng Ji, Michael Auli, Chris Quirk
Citations93
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
Abstract
We introduce Discriminative BLEU (deltaBLEU), a novel metric for intrinsic evaluation of generated text in tasks that admit a diverse range of possible outputs. Reference strings are scored for quality by human raters on a scale of [-1, +1] to weight multi-reference BLEU. In tasks involving generation of conversational responses, deltaBLEU correlates reasonably with human judgments and outperforms sentence-level and IBM BLEU in terms of both Spearman's rho and Kendall's tau.
Keywords
Computer Science
CIDEr: Consensus-based image description evaluation
4,724 Citations2015Ramakrishna Vedantam, C. Lawrence Zitnick +1 more
A novel paradigm for evaluating image descriptions that uses human consensus is proposed and a new automated metric that captures human judgment of consensus better than existing metrics across sentences generated by various sources is evaluated.
METEOR: An Automatic Metric for MT Evaluation with Improved Correlation with Human Judgments
3,723 Citations2005Satanjeev Banerjee, Alon Lavie
METEOR is described, an automatic metric for machine translation evaluation that is based on a generalized concept of unigram matching between the machineproduced translation and human-produced reference translations and can be easily extended to include more advanced matching strategies.
Minimum error rate training in statistical machine translation
2,766 Citations2003Franz Josef Och
It is shown that significantly better results can often be obtained if the final evaluation criterion is taken directly into account as part of the training procedure.
Text REtrieval ConferenceOkapi at TREC
2,229 Citations1994Stephen Robertson, Steve Walker +3 more
Much of the work involved investigating plausible methods of applying Okapi-style weighting to phrases, and expansion using terms from the top documents retrieved by a pilot search on topic terms was used.
Learning deep structured semantic models for web search using clickthrough data
2,044 Citations2013Po-Sen Huang, Xiaodong He +4 more
A series of new latent semantic models with a deep structure that project queries and documents into a common low-dimensional space where the relevance of a document given a query is readily computed as the distance between them are developed.
Automatic evaluation of machine translation quality using n-gram co-occurrence statistics
1,605 Citations2002George R. Doddington
NIST commissioned NIST to develop an MT evaluation facility based on the IBM work, which is now available from NIST and serves as the primary evaluation measure for TIDES MT research.
Journal of Artificial Intelligence ResearchFraming Image Description as a Ranking Task: Data, Models and Evaluation Metrics
1,329 Citations2013Micah Hodosh, Peter Young +1 more
This work proposes to frame sentence-based image annotation as the task of ranking a given pool of captions, and introduces a new benchmark collection, consisting of 8,000 images that are each paired with five different captions which provide clear descriptions of the salient entities and events.
Meteor
1,045 Citations2007Alon Lavie, Abhaya Agarwal
The technical details underlying the Meteor metric are recapped, the latest release includes improved metric parameters and extends the metric to support evaluation of MT output in Spanish, French and German, in addition to English.
National Research Council Canada (Government of Canada)Data-Driven Response Generation in Social Media
576 Citations2011Alan Ritter, Colin Cherry +1 more
It is found that mapping conversational stimuli onto responses is more difficult than translating between languages, due to the wider range of possible responses, the larger fraction of unaligned words/phrases, and the presence of large phrase pairs whose alignment cannot be further decomposed.
Correlating automated and human assessments of machine translation quality
117 Citations2003Deborah A. Coughlin
It is suggested that when human evaluators are forced to make decisions without sufficient context or domain expertise, they fall back on strategies that are not unlike determining n-gram precision.
Accurate Evaluation of Segment-level Machine Translation Metrics
95 Citations2015Yvette Graham, Timothy Baldwin +1 more
Three segment-level metrics — METEOR, NLEPOR and SENTBLEUMOSES — are found to correlate with human assessment at a level not significantly outperformed by any other metric in both the individual language pair assessment for Spanish-toEnglish and the aggregated set of 9 language pairs.
HyTER: Meaning-Equivalent Semantics for Translation Evaluation
72 Citations2012Markus Dreyer, Daniel Marcu
An annotation tool is developed that enables us to create representations that compactly encode an exponential number of correct translations for a sentence, and it is shown that this metric provides better estimates of machine and human translation accuracy than alternative evaluation metrics.
Joint Learning of a Dual SMT System for Paraphrase Generation
72 Citations2012Hong Sun, Ming Zhou
A joint learning method of two SMT systems to optimize the process of paraphrase generation and a revised BLEU score (called iBLEU) which measures the adequacy and diversity of the generated paraphrase sentence is proposed for tuning parameters inSMT systems.
