Joint latent topic models for text and citations
Published 24 August 2008
Ramesh Nallapati, Amr Ahmed, Eric P. Xing, William W. Cohen
Citations413
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
This work addresses the problem of joint modeling of text and citations in the topic modeling framework with two different models called the Pairwise-Link-LDA and the Link-PLSA-Lda models, which combine the LDA and PLSA models into a single graphical model.
Abstract
In this work, we address the problem of joint modeling of text and citations in the topic modeling framework. We present two different models called the Pairwise-Link-LDA and the Link-PLSA-LDA models.
Keywords
Computer Science
Journal of Machine Learning ResearchLatent dirichlet allocation
27,049 Citations2003David M. Blei, Andrew Y. Ng +1 more
Journal of Electronic ImagingPattern Recognition and Machine Learning
21,976 Citations2007Christopher Bishop
Probability Distributions, linear models for Regression, Linear Models for Classification, Neural Networks, Graphical Models, Mixture Models and EM, Sampling Methods, Continuous Latent Variables, Sequential Data are studied.
The PageRank Citation Ranking : Bringing Order to the Web
12,645 Citations1999Lawrence M. Page, Sergey Brin +2 more
This paper describes PageRank, a mathod for rating Web pages objectively and mechanically, effectively measuring the human interest and attention devoted to them, and shows how to efficiently compute PageRank for large numbers of pages.
Journal of the ACMAuthoritative sources in a hyperlinked environment
9,060 Citations1999Jon Kleinberg
This work proposes and test an algorithmic formulation of the notion of authority, based on the relationship between a set of relevant authoritative pages and the set of “hub pages” that join them together in the link structure, and has connections to the eigenvectors of certain matrices associated with the link graph.
A comparison of event models for naive bayes text classification
3,224 Citations1998Andrew McCallum, Kamal Nigam
It is found that the multi-variate Bernoulli performs well with small vocabulary sizes, but that the multinomial performs usually performs even better at larger vocabulary sizes--providing on average a 27% reduction in error over the multi -variateBernoulli model at any vocabulary size.
now publishers, Inc. eBooksGraphical Models, Exponential Families, and Variational Inference
3,154 Citations2007Martin J. Wainwright, Michael I. Jordan
Dynamic topic models
2,322 Citations2006David M. Blei, John Lafferty
A family of probabilistic time series models is developed to analyze the time evolution of topics in large document collections, and dynamic topic models provide a qualitative window into the contents of a large document collection.
arXiv (Cornell University)Probabilistic Latent Semantic Analysis
2,092 Citations2013Thomas Hofmann
This work proposes a widely applicable generalization of maximum likelihood model fitting by tempered EM, based on a mixture decomposition derived from a latent class model which results in a more principled approach which has a solid foundation in statistics.
The link prediction problem for social networks
1,608 Citations2003David Liben‐Nowell, Jon Kleinberg
Experiments on large co-authorship networks suggest that information about future interactions can be extracted from network topology alone, and that fairly subtle measures for detecting node proximity can outperform more direct measures.
Neural Information Processing SystemsCorrelated Topic Models
929 Citations2005John Lafferty, David M. Blei
The correlated topic model (CTM) is developed, where the topic proportions exhibit correlation via the logistic normal distribution and a mean-field variational inference algorithm is derived for approximate posterior inference in this model, which is complicated by the fact that the Logistic normal is not conjugate to the multinomial.
Pachinko allocation
612 Citations2006Wei Li, Andrew McCallum
Improved performance of PAM is shown in document classification, likelihood of held-out data, the ability to support finer-grained topics, and topical keyword coherence.
Proceedings of the National Academy of SciencesMixed-membership models of scientific publications
445 Citations2004Elena A. Erosheva, Stephen E. Fienberg +1 more
This work explores an internal soft-classification structure of articles based only on semantic decompositions of abstracts and bibliographies of PNAS articles, and compares it with the formal discipline classifications.
Link Prediction in Relational Data
442 Citations2003Ben Taskar, Ming-fai Wong +2 more
It is shown that the collective classification approach of RMNs, and the introduction of subgraph patterns over link labels, provide significant improvements in accuracy over flat classification, which attempts to predict each link in isolation.
The Missing Link - A Probabilistic Model of Document Content and Hypertext Connectivity
439 Citations2000David Cohn, Thomas Hofmann
A joint probabilistic model for modeling the contents and inter-connectivity of document collections such as sets of web pages or research paper archives is described, based on a Probabilistic factor decomposition.
The Knowledge Engineering ReviewUncertainty in artificial intelligence
409 Citations1994Simon Parsons
The first conference on Uncertainty in Artificial Intelligence was held in 1985 by a group of people who felt that their views on the use of probability theory were not receiving a fair hearing from the rest of the Al community.
Unsupervised prediction of citation influences
240 Citations2007Laura Dietz, Steffen Bickel +1 more
A probabilistic topic model is devised that explains the generation of documents and incorporates the aspects of topical innovation and topical inheritance via citations, and its ability to predict the strength of influence of citations against manually rated citations is evaluated.
Recommending citations for academic papers
166 Citations2007Trevor Strohman, W. Bruce Croft +1 more
This work uses the text of previous literature as well as the citation graph that connects it to find relevant related material and finds an order of magnitude improvement in mean average precision as compared to a text similarity baseline.
Pachinko allocation: dag-structured mixture models of topic correlations
145 Citations2007Andrew McCallum, Wei Li
Proceedings of the International AAAI Conference on Web and Social MediaLink-PLSA-LDA: A New Unsupervised Model for Topics and Influence of Blogs
125 Citations2021Ramesh Nallapati, William W. Cohen
A new model that can be used to provide a user with highly influential blog postings on the topic of the user's interest is proposed, called Link-PLSA-LDA, that combines PLSA and LDA into a single framework, and explicitly models the topical relationship between the linking and the linked document.
Multiscale topic tomography
59 Citations2007Ramesh Nallapati, Susan Ditmore +2 more
A new probabilistic graphical model is proposed that employs non-homogeneous Poisson processes to model generation of word-counts and its modeling the evolution of topics at various time-scales of resolution, allowing the user to zoom in and out of the time-Scales.
Information genealogy
34 Citations2007Benyah Shaparenko, Thorsten Joachims
This paper proposes a language-modeling approach and a likelihood ratio test to detect influence between documents in a statistically well-founded way and shows how this method can be used to infer citation graphs and to identify the most influential documents in the collection.
