Continuation methods for mixing heterogeneous sources
Generate an AI Snapshot to get a quick, structured summary of this paper.
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
Abstract
The Problem: Many important learning problems involve a compromise between competing sources of information. For example, in semi-supervised classification the available training examples typically include a limited set structure: in estimating graphical models we often need to determine the equivalent sample size, i.e., how the prior model should count relative to the available data. While such “allocation ” problems are ubiquitous, it is not clear how they should be solved in general. In other words, what principle we should use to determine how one privileged source (labeled set, prior) should be balanced relative to the other source (unlabeled set, observed data). In this work we introduce a general method for combining two information sources that continuously visits all possible weightings of the data sources and identifies the one that achieves maximal stability. Motivation: Theoretical homotopy continuation [4] considerations as well as empirical results [1] demonstrate that source allocation problems are not stable in the sense that small changes in source weighting can lead to drastic changes in performance. Because such sudden changes in performance can occur only at a few specific critical weightings of the sources, an estimation algorithm could take advantage of the location of critical weightings to improve its stability. Typical estimation algorithms are unaware of the sensitivity to source weighting, and assume implicitly or explicitly
