login

Data Mining: the search for knowledge in databases.

Data Archiving and Networked Services (DANS)Published 31 January 1994Open access
Marcel Holsheimer, Arno Siebes
Citations184
View PDF

TL;DR

A survey of current data mining research, the main underlying ideas, such as inductive learning, and search strategies and knowledge representations used in data mine systems are presented, and the most important problems and their solutions are described.

Abstract

Data mining is the search for relationships and global patterns that exist in large databases, but are `hidden' among the vast amounts of data, such as a relationship between patient data and their medical diagnosis. These relationships represent valuable knowledge about the database and objects in the database and, if the database is a faithful mirror, of the real world registered by the database. One of the main problems for data mining is that the number of possible relationships is very large, thus prohibiting the search for the correct ones by simple validating each of them. Hence, we need intelligent search strategies, as taken from the area of machine learning. Another important problem is that information in data objects is often corrupted or missing. Hence, statistical techniques should be applied to estimate the reliability of the discovered relationships. This report provides a survey of current data mining research, it presents the main underlying ideas, such as inductive l...

Keywords

Computer Science