On Data and Algorithms: Understanding Inductive Performance
Generate an AI Snapshot to get a quick, structured summary of this paper.
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
This paper addresses two symmetrical issues, the discovery of similarities among classification algorithms, and among datasets, on the basis of error measures, which are used to discover similarities between learners, and both of them to discovering similarities between datasets.
Abstract
Abstract. In this paper we address two symmetrical issues, the discov-ery of similarities among classification algorithms, and among datasets. Both on the basis of error measures, which we use to define the error cor-relation between two algorithms, and determine the relative performance of a list of algorithms. We use the first to discover similarities between learners, and both of them to discover similarities between datasets. The latter sketch maps on the dataset space. Regions within each map exhibit specific patterns of error correlation or relative performance. To acquire an understanding of the factors determining these regions we describe them using simple characteristics of the datasets. Descriptions of each region are given in terms of the distributions of dataset characteristics within it. 1
