login

Comparing pure parallel ensemble creation techniques against bagging

Published 23 April 2004
Lawrence Hall, Kevin W. Bowyer, Robert E. Banfield, Divya Bhadoria, W. Philip Kegelmeyer, Steven A. Eschrich
Citations26

TL;DR

Eight randomization-based approaches to creating an ensemble of decision-tree classifiers are evaluated, and it is found that none of them is consistently more accurate than standard bagging when tested for statistical significance.

Abstract

We experimentally evaluate randomization-based approaches to creating an ensemble of decision-tree classifiers. Unlike methods related to boosting, all of the eight approaches considered here create each classifier in an ensemble independently of the other classifiers. Experiments were performed on 28 publicly available datasets, using C4.5 release 8 as the base classifier. While each of the other seven approaches has some strengths, we find that none of them is consistently more accurate than standard bagging when tested for statistical significance.

Keywords

EngineeringBusiness, Management and Accounting