login

Building a Cross-Language Entity Linking Collection in Twenty-One Languages

Lecture notes in computer sciencePublished 1 January 2011
James Mayfield, Dawn Lawrie, Paul McNamee, Douglas W. Oard
Citations17
SJR quartileQ2
SJR score0.35
SNIP0.55

TL;DR

An efficient way to create a test collection for evaluating the accuracy of cross-language entity linking is described, which includes approximately 55,000 queries, comprising between 875 and 4,329 queries for each of twenty-one non-English languages.

Abstract

We describe an efficient way to create a test collection for evaluating the accuracy of cross-language entity linking. Queries are created by semi-automatically identifying person names on the English side of a parallel corpus, using judgments obtained through crowdsourcing to identify the entity corresponding to the name, and projecting the English name onto the non-English document using word alignments. We applied the technique to produce the first publicly available multilingual cross-language entity linking collection. The collection includes approximately 55,000 queries, comprising between 875 and 4,329 queries for each of twenty-one non-English languages.

Keywords

Computer ScienceDecision Sciences