login

Identifying Co-referential Names Across Large Corpora

Lecture notes in computer sciencePublished 1 January 2006
Levon Lloyd, Andrew Mehler, Steven Skiena
Citations17
SJR quartileQ2
SJR score0.35
SNIP0.55

TL;DR

This work presents an algorithm for finding all co-reference sets in a large corpus of real news text by involves three steps: morphological similarity detection, contextual similarity analysis, and clustering.

Abstract

A single logical entity can be referred to by several different names over a large text corpus. We present our algorithm for finding all such co-reference sets in a large corpus. Our algorithm involves three steps: morphological similarity detection, contextual similarity analysis, and clustering. Finally, we present experimental results on over large corpus of real news text to analyze the performance our techniques.

Keywords

Computer ScienceDecision Sciences