login

Proceedings of the NAACL HLT 2010 Workshop on Creating Speech and Language Data with Amazon's Mechanical Turk

Published 6 June 2010
Chris Callison-Burch, Mark Dredze
Citations141

TL;DR

The NAACL-2010Workshop on Creating Speech and Language Data with Amazon's Mechanical Turk explores applications of crowdsourcing technologies for the creation and study of language data and investigates new ways of integrating user feedback in the learning process.

Abstract

The NAACL-2010Workshop on Creating Speech and Language DataWith Amazon's Mechanical Turk explores applications of crowdsourcing technologies for the creation and study of language data. Recent work has evaluated the effectiveness of using crowdsourcing platforms, such as Amazon's Mechanical Turk, to create annotated data for natural language processing applications. This workshop further explores this area and these proceedings contain 34 papers and an overview paper that each experiment with applications of Mechanical Turk. The diversity of applications showcases the new possibilities for annotating speech and text, and has the potential to dramatically change how we create data for human language technologies. Papers in the workshop also looked at best practices in creating data using Mechanical Turk. Experiments evaluated how to design Human Intelligence Tasks (HITs), how to attract users to the task, how to price annotation tasks, and how to ensure data quality. Applications include the creation of data sets for standard NLP tasks, developing entirely new tasks, and investigating new ways of integrating user feedback in the learning process. The workshop featured an open-ended shared task in which 35 teams were awarded $100 of credit on Amazon Mechanical Turk to spend on an annotation task of their choosing. Results of the shared task are described in short papers and all collected data is publicly available. Shared task participants focused on data collection questions, such as how to convey complex tasks to non-experts, how to evaluate and ensure quality and annotation cost and speed.

Keywords

Computer Science