Mining entity-identification rules for database integration
Generate an AI Snapshot to get a quick, structured summary of this paper.
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
This work proposes the use of distances between attribute values as a measure of similarity between the records they represent, and shows how knowledge discovery techniques can be used to automatically derive conditions for EI directly from the data, using a distance-based framework.
Abstract
Entity identification (EI) is the identification and in-tegration of all records which represent he same real-world entity, and is an important task in database integration process. When a common identification mechanism for similar records across heterogeneous databases is not readily available, EI is performed by examining the relationships between various attribute values among the records. We propose the use of dis-tances between attribute values as a measure of sim-ilarity between the records they represent. Record-matching conditions for EI can then be expressed as constraints on the attribute distances. We show how knowledge discovery techniques can be used to auto-matically derive these conditions (expressed as deci-sion trees) directly from the data, using a distance-based framework.
