An event-centric model for multilingual document similarity
Generate an AI Snapshot to get a quick, structured summary of this paper.
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
A novel document similarity measure that is based on events extracted from documents that is term- and language-independent as temporal and geographic expressions mentioned in texts are normalized to a standard format and allows to determine similar documents across languages.
Abstract
Document similarity measures play an important role in many document retrieval and exploration tasks. Over the past decades, several models and techniques have been developed to determine a ranked list of documents similar to a given query document. Interestingly, the proposed approaches typically rely on extensions to the vector space model and are rarely suited for multilingual corpora.
