login

Distributed Inference for Latent Dirichlet Allocation

Published 3 December 2007
David Newman, Padhraic Smyth, Max Welling, Arthur Asuncion
Citations220

TL;DR

Using five real-world text corpora, it is shown that distributed learning works very well for LDA models, i.e., perplexity and precision-recall scores for distributed learning are indistinguishable from those obtained with single-processor learning.

Abstract

1 Introduction Very large data sets, such as collections of images, text, and related data, are becoming increasinglycommon, with examples ranging from digitized collections of books by companies such as Google and Amazon, to large collections of images at Web sites such as Flickr, to the recent Netflix customerrecommendation data set. These data sets present major opportunities for machine learning, such as the ability to explore much richer and more expressive models, as well as providing new andinteresting domains for the application of learning algorithms.

Keywords

Computer Science