login

A correlated topic model of Science

The Annals of Applied StatisticsPublished 1 June 2007Open access
David M. Blei, John D. Lafferty
Citations921
SJR quartileQ1
SJR score0.94
SNIP0.93
View PDF

TL;DR

The correlated topic model (CTM) is developed, where the topic proportions exhibit correlation via the logistic normal distribution, and it is demonstrated its use as an exploratory tool of large document collections.

Abstract

Topic models, such as latent Dirichlet allocation (LDA), can be\nuseful tools for the statistical analysis of document\ncollections and other discrete data. The LDA model assumes that\nthe words of each document arise from a mixture of\ntopics, each of which is a distribution over the\nvocabulary. A limitation of LDA is the inability to model topic\ncorrelation even though, for example, a document about genetics\nis more likely to also be about disease than X-ray astronomy.\nThis limitation stems from the use of the Dirichlet distribution\nto model the variability among the topic proportions. In this\npaper we develop the correlated topic model (CTM), where the\ntopic proportions exhibit correlation via the logistic normal\ndistribution [J. Roy. Statist. Soc. Ser. B\n44 (1982) 139–177]. We derive a fast variational\ninference algorithm for approximate posterior inference in this\nmodel, which is complicated by the fact that the logistic normal\nis not conjugate to the multinomial. We apply the CTM to the\narticles from Science published from 1990–1999, a data\nset that comprises 57M words. The CTM gives a better fit of the\ndata than LDA, and we demonstrate its use as an exploratory tool\nof large document collections.

Keywords

Computer ScienceBiochemistry, Genetics and Molecular BiologySocial Sciences