Information Retrieval Meta-Evaluation: Challenges And Opportunities In The Music Domain.
Published 24 October 2011Open access
Julián Urbano
Citations23
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
A survey of past meta-evaluation work in the context of Text Information Retrieval argues that the music community still needs to address various issues concerning the evaluation of music systems and the IR cycle, pointing out directions for further research and proposals in this line.
Abstract
[TODO] Add abstract here.
Keywords
Computer ScienceArts and Humanities
ACM Transactions on Information SystemsCumulated gain-based evaluation of IR techniques
4,634 Citations2002Kalervo Järvelin, Jaana Kekäläinen
This article proposes several novel measures that compute the cumulative gain the user obtains by examining the retrieval result up to a given ranked position, and test results indicate that the proposed measures credit IR methods for their ability to retrieve highly relevant documents and allow testing of statistical significance of effectiveness differences.
Columbia Academic Commons (Columbia University)The Million Song Dataset
829 Citations2011Thierry Bertin-Mahieux, Daniel P. W. Ellis +2 more
The Million Song Dataset, a freely-available collection of audio features and metadata for a million contemporary popular music tracks, is introduced and positive results on year prediction are shown, and the future development of the dataset is discussed.
Retrieval evaluation with incomplete information
733 Citations2004Chris Buckley, Ellen M. Voorhees
It is shown that current evaluation measures are not robust to substantially incomplete relevance judgments, and a new measure is introduced that is both highly correlated with existing measures when complete judgments are available and more robust to incomplete judgment sets.
A comparison of statistical significance tests for information retrieval evaluation
717 Citations2007Mark D. Smucker, James Allan +1 more
It is discovered that there is little practical difference between the randomization, bootstrap, and t tests and their use should be discontinued for measuring the significance of a difference between means.
How reliable are the results of large-scale information retrieval experiments?
577 Citations1998Justin Zobel
A detailed empirical investigation of the TREC results shows that the measured relative performance of systems appears to be reliable, but that recall is overestimated: it is likely that many relevant documents have not been found.
ACM SIGIR ForumEvaluating Evaluation Measure Stability
567 Citations2017Chris Buckley, Ellen M. Voorhees
A novel way of examining the accuracy of the evaluation measures commonly used in information retrieval experiments is presented, which validates several of the rules-of-thumb experimenters use and challenges other beliefs, such as the common evaluation measures are equally reliable.
Information Processing & ManagementVariations in relevance judgments and the measurement of retrieval effectiveness
505 Citations2000Ellen M. Voorhees
Lecture notes in computer scienceThe Philosophy of Information Retrieval Evaluation
385 Citations2002Ellen M. Voorhees
The fundamental assumptions and appropriate uses of the Cranfield paradigm, especially as they apply in the context of the evaluation conferences, are reviewed.
Information retrieval system evaluation
365 Citations2005Mark Sanderson, Justin Zobel
It is found that the t-test is highly reliable (more so than the sign or Wilcoxon test), and is far more reliable than simply showing a large percentage difference in effectiveness measures between IR systems.
Estimating average precision with incomplete and imperfect judgments
329 Citations2006Emine Yılmaz, Javed A. Aslam
This work proposes three evaluation measures that are approximations to average precision even when the relevance judgments are incomplete and are more robust to incomplete or imperfect relevance judgments than bpref, and proposes estimates of average precision that are simple and accurate.
The effect of topic set size on retrieval experiment error
246 Citations2002Ellen M. Voorhees, Chris Buckley
Using TREC results to empirically derive error rates based on the number of topics used in a test and the observed difference in the average scores indicates researchers need to take care when concluding one method is better than another, especially if few topics are used.
Minimal test collections for retrieval evaluation
242 Citations2006Ben Carterette, James Allan +1 more
This work links evaluation with test collection construction to gain an understanding of the minimal judging effort that must be done to have high confidence in the outcome of an evaluation.
Improvements that don't add up
238 Citations2009Timothy G. Armstrong, Alistair Moffat +2 more
This paper analyzes results achieved on the TREC Ad-Hoc, Web, Terabyte, and Robust collections as reported in SIGIR and CIKM and proposes a practice of regular longitudinal comparison to ensure measurable progress, or at least prevent the lack of it from going unnoticed.
Ranking retrieval systems without relevance judgments
236 Citations2001Ian Soboroff, Charles Nicholas +1 more
The initial results of a new evaluation methodology which replaces human relevance judgments with a randomly selected mapping of documents to topics are proposed, which are referred to aspseudo-relevance judgments.
A simple and efficient sampling method for estimating AP and NDCG
232 Citations2008Emine Yılmaz, Evangelos Kanoulas +1 more
Evaluation by highly relevant documents
228 Citations2001Ellen M. Voorhees
To explore the role highly relevant documents play in retrieval system evaluation, assessors for the TREC-9 web track used a three-point relevance scale and also selected best pages for each topic, confirming the hypothesis that different retrieval techniques work better for retrievinghighly relevant documents.
The significance of the Cranfield tests on index languages
181 Citations1991Cyril W. Cleverdon
Relevance assessment
177 Citations2008Peter Bailey, Nick Craswell +4 more
It appears that test collections are not completely robust to changes of judge when these judges vary widely in task and topic expertise, and both system scores and system rankings are subject to consistent but small differences across the three assessment sets.
Do user preferences and evaluation measures line up?
159 Citations2010Mark Sanderson, Monica Lestari Paramita +2 more
It is established that preferences and evaluation measures correlate: systems measured as better on a test collection are preferred by users and the nDCG measure is found to correlate best with user preferences compared to a selection of other well known measures.
Why batch and user evaluations do not give the same results
156 Citations2001Andrew Turpin, William Hersh
Assessment of the TREC Interactive Track showed that while the queries entered by real users into systems yielding better results in batch studies gave comparable gains in ranking of relevant documents for those users, they did not translate into better performance on specific tasks.
The 2007 MIREX Audio Mood Classification Task: Lessons Learned
144 Citations2008Xiao Hu, J. Stephen Downie +3 more
Important issues in setting up the AMC task are described, dataset construction and ground-truth labeling are analyzed, and human assessments on the audio dataset, as well as system performances from various angles are analyzed.
Information RetrievalOn information retrieval metrics designed for evaluation with incomplete relevance assessments
128 Citations2008Tetsuya Sakai, Noriko Kando
This article compares the robustness of IR metrics to incomplete relevance assessments, using four different sets of graded-relevance test collections with submitted runs—the TREC 2003 and 2004 robust track data and the NTCIR-6 Japanese and Chinese IR data from the crosslingual task.
Information RetrievalBias and the limits of pooling for large collections
117 Citations2007Chris Buckley, Darrin L. Dimmick +2 more
It is shown that the judgment sets produced by traditional pooling when the pools are too small relative to the total document set size can be biased in that they favor relevant documents that contain topic title words.
Studies in computational intelligenceThe Music Information Retrieval Evaluation eXchange: Some Observations and Insights
110 Citations2010J. Stephen Downie, Andreas F. Ehmann +2 more
This chapter outlines some of the major highlights of the past four years of MIREX evaluations, including its organizing principles, the selection of evaluation metrics, and the evolution of evaluation tasks.
Teaching SociologyEvaluating Information: A Guide for Users of Social Science Research
102 Citations1999Jacqueline Carrigan, Jeffrey Katzer +2 more
Information Processing & ManagementBinary and graded relevance in IR evaluations—Comparison of the effects on ranking of IR systems
98 Citations2005Jaana Kekäläinen
In this study the rankings of IR systems based on binary and graded relevance in TREC 7 and 8 data are compared and the results show the different character of the measures.
The effect of assessor error on IR system evaluation
90 Citations2010Ben Carterette, Ian Soboroff
This paper examines the robustness of the TREC Million Query track methods when some assessors make significant and systematic errors, and finds that while averages are robust, assessor errors can have a large effect on system rankings.
Computer Music JournalThe Scientific Evaluation of Music Information Retrieval Systems: Foundations and Future
82 Citations2004J. Stephen Downie
This article provides an overview of the current scientific problem facing MIR research and reports upon the findings of the Music Information Retrieval (MIR)/ Music Digital Library (MDL) Evaluation Frameworks Project.
Statistical power in retrieval experimentation
74 Citations2008William Webber, Alistair Moffat +1 more
It is demonstrated that greater statistical power is achieved for the same relevance assessment effort by evaluating a large number of topics shallowly than a small number deeply, and that it leads to a bias in favour of finding both power and significance.
Robust test collections for retrieval evaluation
58 Citations2007Ben Carterette
This work formally defines what it means for judgments to be reusable: the confidence in an evaluation of new systems can be accurately assessed from an existing set of relevance judgments, and presents a method for augmenting a set ofrelevant judgments with relevance estimates that require no additional assessor effort.
ACM Transactions on Information SystemsA few good topics
57 Citations2009John Guiver, Stefano Mizzaro +1 more
Using a variety of effectiveness metrics and measures of goodness of prediction, a study of a set of TREC and NTCIR results confirms the hypothesis that some topics or topic sets are better than others at predicting true system effectiveness, and provides evidence that the value of a topic set for this purpose does generalize.
Zenodo (CERN European Organization for Nuclear Research)Crowdsourcing Music Similarity Judgments Using Mechanical Turk.
48 Citations2010Jin Ha Lee
This paper compares the similarity judgments collected from Evalutron6000 (E6K) and MTurk using the Music Information Retrieval Evaluation eXchange 2009 Audio Music Similarity and Retrival task dataset, and concludes that using M Turk is a practical alternative of music similarity when it is used with some precautions.
ACM SIGIR ForumCrowdsourcing for search evaluation
47 Citations2011Vitor R. Carvalho, Matthew Lease +1 more
The Crowdsourcing for Search Evaluation Workshop (CSE 2010) was held on July 23, 2010 in Geneva, Switzerland, in conjunction with the 33rd Annual ACM SIGIR Conference.
Data Archiving and Networked Services (DANS)A Ground Truth For Half A Million Musical Incipits
45 Citations2005Rainer Typke, Marc den Hoed +3 more
This work filtered the RISM A/II collection so that about 50 candidates per query were left, which was then presented to 35 human experts for a final ranking, and obtained ground truths, which can be used for evaluating music information retrieval systems.
Crowdsourcing Preference Judgments for Evaluation of Music Similarity Tasks
41 Citations2010Julián Urbano, Jorge Morato +2 more
It is shown that crowdsourcing is a perfectly viable alternative to evaluate music systems without the need for experts, and produces lists very similar to the original ones, while dealing with some defects of the original methodology.
A Measure for Evaluating Retrieval Techniques based on Partially Ordered Ground Truth Lists
37 Citations2006Rainer Typke, Remco C. Veltkamp +1 more
The "average dynamic recall" (ADR) is introduced that averages the recall among a dynamic set of relevant documents, taking into account the fact that the ground truth reliably orders groups of matches, but not always individual matches.
Precision-at-ten considered redundant
34 Citations2008William Webber, Alistair Moffat +2 more
It is demonstrated that complex metrics are as good as or better than simple metrics at predicting the performance of the simple metrics on other topics.
International Symposium/Conference on Music Information RetrievalHuman Similarity Judgments: Implications For The Design Of Formal Evaluations.
32 Citations2007M. Cameron Jones, J. Stephen Downie +1 more
Findings of a series of analyses of human similarity judgments from the Symbolic Melodic Similarity, and Audio Music Similarity tasks from the Music Information Retrieval Evaluation Exchange (MIREX) 2006 are presented.
Lecture notes in computer scienceIf I Had a Million Queries
31 Citations2009Ben Carterette, Virgil Pavlu +3 more
This work analyzes a sample of queries categorized by length and corpus-appropriateness to determine the right proportion needed to distinguish between systems, and shows that only a small, biased sample of query with sparse judgments is needed to produce the same results as a much larger sample of questions.
Measuring the reusability of test collections
26 Citations2010Ben Carterette, Evgeniy Gabrilovich +2 more
The proposed methods provide simple yet highly effective tests for determining whether an existing set of judgments is useful for evaluating a new system and can reliably estimate confidence intervals that are indicative of collection reusability.
Reusable test collections through experimental design
23 Citations2010Ben Carterette, Evangelos Kanoulas +2 more
This work proposes a methodology for information retrieval experimentation that collects evidence for or against the reusability of a test collection while judgments are being made, and provides a description of an actual implementation of the framework for creating a large test collection.
Zenodo (CERN European Organization for Nuclear Research)Audio Cover Song Identification: Mirex 2006-2007 Results And Analyses.
23 Citations2008J. Stephen Downie, Mert Bay +2 more
Analysis of the 2006 and 2007 results of the Music Information Retrieval Evaluation eXchange (MIREX) Audio Cover Song Identification (ACS) tasks indicate significant improvements in this domain have been made over the course of 2006-2007.
Musiclef: A Benchmark Activity In Multimodal Music Information Retrieval.
21 Citations2011Nicola Orio, David Rizo +4 more
This work presents the rationale, tasks and procedures of MusiCLEF, a novel benchmarking activity that has been developed along with the Cross-Language Evaluation Forum (CLEF), to promote the development of new methodologies for music access and retrieval on real public music collections.
Lecture notes in computer scienceOn the Contributions of Topics to System Evaluation
20 Citations2011Stephen Robertson
It turns out to be hard to establish generalisability; it is not at all clear that it is possible to identify subsets of topics that are good for general evaluation.
Arrow@dit (Dublin Institute of Technology)Evaluating Evaluation Measures
19 Citations2007Ines Rehbein, Josef van Genabith
An analysis of specic error types indicates that the dependency-based evaluation is most appropriate to reect parse quality, and shows that PARSEVAL should not be used to compare parser performance for parsers trained on treebanks with different annotation schemes.
Improving The Generation Of Ground Truths Based On Partially Ordered Lists.
14 Citations2010Julián Urbano, Mónica Marrero +2 more
It is shown that it is not possible to ensure lists completely consistent, and a measure of consistency based on Average Dynamic Recall is developed and several alternatives to arrange the lists prove to be more consistent than the original method.
Audio Music Similarity And Retrieval: Evaluation Power And Stability.
11 Citations2011Julián Urbano, Diego Martín +2 more
The reliability of the res ults in the evaluation of Audio Music Similarity and Retrieval systems is analyzed and it is concluded that experimenters can be very confident that if a significant difference is found between systems, the difference is indeed real.
ACM SIGIR ForumBeyond binary relevance
10 Citations2008Paul N. Bennett, Ben Carterette +2 more
The goal of the workshop was to examine how the type of response elicited from assessors and users influences the evaluation and analysis of retrieval and filtering applications and explored research challenges at the intersection of novel measures of relevance, novel learning methods, and core evaluation issues.
VTechWorks (Virginia Tech)Information Retrieval System Evaluation
8 Citations2012Shiyi Wei, Victoria Suwardiman +1 more
This module introduces the evaluation in information retrieval by focusing on the standard measurement of system effectiveness through relevance judgments, and requires the Lucidworks software to perform the exercises.
On The Importance Of "Real" Audio Data For Mir Algorithm Evaluation At The Note-Level - A Comparative Study.
5 Citations2011Bernhard Niedermayer, Sebastian Böck +1 more
It is shown that the algorithms’ performance on the different datasets varies considerably, but synthesized audio, does not necessarily yield better results.
