login

Segmentation of Software Engineering Datasets Using the M5 Algorithm

Lecture notes in computer sciencePublished 1 January 2006Open access
Daniel Rodríguez, J. J. Cuadrado, Miguel‐Ángel Sicilia, Roberto Ruíz
Citations8
SJR quartileQ2
SJR score0.35
SNIP0.55
View PDF

TL;DR

An empirical study that uses clustering techniques to derive segmented models from software engineering repositories and shows that there is an improvement in the accuracy of the results when using clustering.

Abstract

This paper reports an empirical study that uses clustering techniques to derive segmented models from software engineering repositories, focusing on the improvement of the accuracy of estimates. In particular, we used two datasets obtained from the International Software Benchmarking Standards Group (ISBSG) repository and created clusters using the M5 algorithm. Each cluster is associated with a linear model. We then compare the accuracy of the estimates so generated with the classical multivariate linear regression and least median squares. Results show that there is an improvement in the accuracy of the results when using clustering. Furthermore, these techniques can help us to understand the datasets better; such techniques provide some advantages to project managers while keeping the estimation process within reasonable complexity.

Keywords

Computer Science