Building a Cross-Language Entity Linking Collection in Twenty-One Languages
Generate an AI Snapshot to get a quick, structured summary of this paper.
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
An efficient way to create a test collection for evaluating the accuracy of cross-language entity linking is described, which includes approximately 55,000 queries, comprising between 875 and 4,329 queries for each of twenty-one non-English languages.
Abstract
We describe an efficient way to create a test collection for evaluating the accuracy of cross-language entity linking. Queries are created by semi-automatically identifying person names on the English side of a parallel corpus, using judgments obtained through crowdsourcing to identify the entity corresponding to the name, and projecting the English name onto the non-English document using word alignments. We applied the technique to produce the first publicly available multilingual cross-language entity linking collection. The collection includes approximately 55,000 queries, comprising between 875 and 4,329 queries for each of twenty-one non-English languages.
