login

Continuation methods for mixing heterogeneous sources

Published 1 August 2002
Adrian Corduneanu, Tommi Jaakkola
Citations31

Abstract

The Problem: Many important learning problems involve a compromise between competing sources of information. For example, in semi-supervised classification the available training examples typically include a limited set structure: in estimating graphical models we often need to determine the equivalent sample size, i.e., how the prior model should count relative to the available data. While such “allocation ” problems are ubiquitous, it is not clear how they should be solved in general. In other words, what principle we should use to determine how one privileged source (labeled set, prior) should be balanced relative to the other source (unlabeled set, observed data). In this work we introduce a general method for combining two information sources that continuously visits all possible weightings of the data sources and identifies the one that achieves maximal stability. Motivation: Theoretical homotopy continuation [4] considerations as well as empirical results [1] demonstrate that source allocation problems are not stable in the sense that small changes in source weighting can lead to drastic changes in performance. Because such sudden changes in performance can occur only at a few specific critical weightings of the sources, an estimation algorithm could take advantage of the location of critical weightings to improve its stability. Typical estimation algorithms are unaware of the sensitivity to source weighting, and assume implicitly or explicitly

Keywords

Computer ScienceDecision Sciences