Identifying Co-referential Names Across Large Corpora
Generate an AI Snapshot to get a quick, structured summary of this paper.
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
This work presents an algorithm for finding all co-reference sets in a large corpus of real news text by involves three steps: morphological similarity detection, contextual similarity analysis, and clustering.
Abstract
A single logical entity can be referred to by several different names over a large text corpus. We present our algorithm for finding all such co-reference sets in a large corpus. Our algorithm involves three steps: morphological similarity detection, contextual similarity analysis, and clustering. Finally, we present experimental results on over large corpus of real news text to analyze the performance our techniques.
