A correlated topic model of Science
Generate an AI Snapshot to get a quick, structured summary of this paper.
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
The correlated topic model (CTM) is developed, where the topic proportions exhibit correlation via the logistic normal distribution, and it is demonstrated its use as an exploratory tool of large document collections.
Abstract
Topic models, such as latent Dirichlet allocation (LDA), can be\nuseful tools for the statistical analysis of document\ncollections and other discrete data. The LDA model assumes that\nthe words of each document arise from a mixture of\ntopics, each of which is a distribution over the\nvocabulary. A limitation of LDA is the inability to model topic\ncorrelation even though, for example, a document about genetics\nis more likely to also be about disease than X-ray astronomy.\nThis limitation stems from the use of the Dirichlet distribution\nto model the variability among the topic proportions. In this\npaper we develop the correlated topic model (CTM), where the\ntopic proportions exhibit correlation via the logistic normal\ndistribution [J. Roy. Statist. Soc. Ser. B\n44 (1982) 139–177]. We derive a fast variational\ninference algorithm for approximate posterior inference in this\nmodel, which is complicated by the fact that the logistic normal\nis not conjugate to the multinomial. We apply the CTM to the\narticles from Science published from 1990–1999, a data\nset that comprises 57M words. The CTM gives a better fit of the\ndata than LDA, and we demonstrate its use as an exploratory tool\nof large document collections.
