A universal theorem on learning curves
Generate an AI Snapshot to get a quick, structured summary of this paper.
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
A universal asymptotic behavior of learning curves for general noiseless dichotomy machines, or neural networks is proved, it is proved that irrespective of the architecture of a machine, the average predictive entropy or the information gain converges to 0.
Abstract
A learning curve shows how fast a learning machine improves its behavior as the number of training examples increases. This paper proves a universal asymptotic behavior of learning curves for general noiseless dichotomy machines, or neural networks. It is proved that irrespective of the architecture of a machine, the average predictive entropy or the information gain 〈e∗(t)〉 converges to 0 as 〈e∗(t)〉 ∼ d/t as the number t of training exampies increases, where d is the number of modifiable parameters of a machine.
