login

Mining entity-identification rules for database integration

Published 2 August 1996
M. Ganesh, Jaideep Srivastava, Travis D. Richardson
Citations21

TL;DR

This work proposes the use of distances between attribute values as a measure of similarity between the records they represent, and shows how knowledge discovery techniques can be used to automatically derive conditions for EI directly from the data, using a distance-based framework.

Abstract

Entity identification (EI) is the identification and in-tegration of all records which represent he same real-world entity, and is an important task in database integration process. When a common identification mechanism for similar records across heterogeneous databases is not readily available, EI is performed by examining the relationships between various attribute values among the records. We propose the use of dis-tances between attribute values as a measure of sim-ilarity between the records they represent. Record-matching conditions for EI can then be expressed as constraints on the attribute distances. We show how knowledge discovery techniques can be used to auto-matically derive these conditions (expressed as deci-sion trees) directly from the data, using a distance-based framework.

Keywords

Computer Science