login

Mining Complex Data Streams: Discretization, Attribute Selection and Classification

Journal of Advances in Information TechnologyPublished 1 August 2013Open access
Dewan Md. Farid, Chowdhury Mofizur Rahman
Citations42
SJR quartileQ3
SJR score0.30
SNIP0.63
View PDF

TL;DR

The proposed discretization algorithm finds the possible cut points in continuous attributes using information gain heuristic and naive Bayesian classifier that can separate the class distributions.

Abstract

Due to the large volume of data set as well as complex and dynamic properties of data instances, several data mining algorithms have been applied for mining complex data streams in the last decades. Now a day, knowledge extraction from data streams is getting more complex because the structure of the data instance does not match the attribute values when considering the tabulated data, texts, web, images or videos etc. In this paper, we address some difficulties of mining complex data streams such as dealing with continuous attributes, input attribute selection, and classifier construction. The proposed discretization algorithm finds the possible cut points in continuous attributes using information gain heuristic and naïve Bayesian classifier that can separate the class distributions. We evaluate the proposed algorithms on several benchmark data sets from UCI machine learning repository. The experimental results demonstrate that the proposed method improves the quality of discretization of continuous attributes and scales up the classification accuracy for different types of classification problem.

Keywords

Computer Science