Mining Text Data
Published 1 January 2015
Charų C. Aggarwal
Citations614
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
Abstract
Text data are copiously found in many domains, such as the Web, social networks, newswire services, and libraries. With the increasing ease in archival of human speech and expression, the volume of text data will only increase over time. This trend is reinforced by the increasing digitization of libraries and the ubiquity of the Web and social networks.
Keywords
Computer Science
Journal of the Royal Statistical Society Series B (Statistical Methodology)Maximum Likelihood from Incomplete Data Via the <i>EM</i> Algorithm
49,657 Citations1977A. P. Dempster, N. M. Laird +1 more
The Nature of Statistical Learning Theory
39,279 Citations1995Vladimir Vapnik
Elements of Information Theory
37,533 Citations2001Thomas M. Cover, Joy A. Thomas
Machine LearningSupport-Vector Networks
33,035 Citations1995Corinna Cortes, Vladimir Vapnik
High generalization ability of support-vector networks utilizing polynomial input transformations is demonstrated and the performance of the support- vector network is compared to various classical learning algorithms that all took part in a benchmark study of Optical Character Recognition.
Journal of Machine Learning ResearchLatent dirichlet allocation
27,049 Citations2003David M. Blei, Andrew Y. Ng +1 more
TechnometricsStatistical Learning Theory
26,913 Citations1999Yuhai Wu, Vladimir Vapnik
Presenting a method for determining the necessary and sufficient conditions for consistency of learning process, the author covers function estimates from small data pools, applying these estimations to real-life problems, and much more.
Proceedings of the IEEEA tutorial on hidden Markov models and selected applications in speech recognition
22,785 Citations1989L. R. Rabiner
Journal of Electronic ImagingPattern Recognition and Machine Learning
21,976 Citations2007Christopher Bishop
Probability Distributions, linear models for Regression, Linear Models for Classification, Neural Networks, Graphical Models, Mixture Models and EM, Sampling Methods, Continuous Latent Variables, Sequential Data are studied.
Journal of Computer and System SciencesA Decision-Theoretic Generalization of On-Line Learning and an Application to Boosting
20,320 Citations1997Yoav Freund, Robert E. Schapire
IEEE Transactions on Pattern Analysis and Machine IntelligenceNormalized cuts and image segmentation
15,696 Citations2000Jianbo Shi, Jitendra Malik
Encyclopedia of Statistics in Behavioral SciencePrincipal Component Analysis
14,549 Citations2005Ian T. Jolliffe
NatureLearning the parts of objects by non-negative matrix factorization
14,055 Citations1999Daniel D. Lee, H. Sebastian Seung
An algorithm for non-negative matrix factorization is demonstrated that is able to learn parts of faces and semantic features of text and is in contrast to other methods that learn holistic, not parts-based, representations.
ScholarlyCommons (University of Pennsylvania)Conditional Random Fields: Probabilistic Models for Segmenting and Labeling Sequence Data
12,978 Citations2001John Lafferty, Andrew McCallum +1 more
This work presents iterative parameter estimation algorithms for conditional random fields and compares the performance of the resulting models to HMMs and MEMMs on synthetic and natural-language data.
Journal of the American Society for Information ScienceIndexing by latent semantic analysis
12,677 Citations1990Scott Deerwester, Susan Dumais +3 more
The MIT Press eBooksGaussian Processes for Machine Learning
10,409 Citations2005Carl Edward Rasmussen, Christopher K. I. Williams
The book provides a long-needed, systematic and unified treatment of theoretical and practical aspects of GPs in machine learning, targeted at researchers and students in machine learning and applied statistics.
Statistics and ComputingA tutorial on spectral clustering
10,269 Citations2007Ulrike von Luxburg
This tutorial describes different graph Laplacians and their basic properties, present the most common spectral clustering algorithms, and derive those algorithms from scratch by several different approaches.
The MIT Press eBooksLearning with Kernels
9,548 Citations2001Bernhard Schölkopf, Alexander J. Smola
Learning with Kernels provides an introduction to SVMs and related kernel methods that provide all of the concepts necessary to enable a reader equipped with some basic mathematical knowledge to enter the world of machine learning using theoretically well-founded yet easy-to-use kernel algorithms.
Information Processing & ManagementTerm-weighting approaches in automatic text retrieval
9,532 Citations1988Gerard Salton, Chris Buckley
This paper summarizes the insights gained in automatic term weighting, and provides baseline single term indexing models with which other more elaborate content analysis procedures can be compared.
Lecture notes in computer scienceText categorization with Support Vector Machines: Learning with many relevant features
7,925 Citations1998Thorsten Joachims
SVMs achieve substantial improvements over the currently best performing methods and behave robustly over a variety of di-erent learning tasks, eliminating the need for manual parameter tuning.
ACM Computing SurveysMachine learning in automated text categorization
7,899 Citations2002Fabrizio Sebastiani
This survey discusses the main approaches to text categorization that fall within the machine learning paradigm and discusses in detail issues pertaining to three different problems, namely, document representation, classifier construction, and classifier evaluation.
SIAM Journal on OptimizationA Singular Value Thresholding Algorithm for Matrix Completion
5,870 Citations2010Jian‐Feng Cai, Emmanuel J. Candès +1 more
This paper develops a simple first-order and easy-to-implement algorithm that is extremely efficient at addressing problems in which the optimal solution has low rank, and develops a framework in which one can understand these algorithms in terms of well-known Lagrange multiplier algorithms.
Proceedings of the IEEEThe viterbi algorithm
5,595 Citations1973G. David Forney
This paper gives a tutorial exposition of the Viterbi algorithm and of how it is implemented and analyzed, and increasing use of the algorithm in a widening variety of areas is foreseen.
Freebase
4,892 Citations2008Kurt Bollacker, Colin Evans +3 more
MQL provides an easy-to-use object-oriented interface to the tuple data in Freebase and is designed to facilitate the creation of collaborative, Web-based data-oriented applications.
IEEE ASSP MagazineAn introduction to hidden Markov models
4,789 Citations1986L. R. Rabiner, Biing‐Hwang Juang
The purpose of this tutorial paper is to give an introduction to the theory of Markov models, and to illustrate how they have been applied to problems in speech recognition.
Journal of DocumentationA STATISTICAL INTERPRETATION OF TERM SPECIFICITY AND ITS APPLICATION IN RETRIEVAL
4,442 Citations1972Karen Spärck Jones
It is argued that terms should be weighted according to collection frequency, so that matches on less frequent, more specific, terms are of greater value than matches on frequent terms.
International Journal of LexicographyIntroduction to WordNet: An On-line Lexical Database<sup>*</sup>
4,207 Citations1990George A. Miller, Richard Beckwith +3 more
Standard alphabetical procedures for organizing lexical information put together words that are spelled alike and scatter words with similar or related meanings haphazardly through the list.
BIRCH
3,906 Citations1996Tian Zhang, Raghu Ramakrishnan +1 more
A data clustering method named BIRCH (Balanced Iterative Reducing and Clustering using Hierarchies) is presented, and it is demonstrated that it is especially suitable for very large databases.
Journal of the American Statistical AssociationHierarchical Dirichlet Processes
3,567 Citations2006Yee Whye Teh, Michael I. Jordan +2 more
This work considers problems involving groups of data where each observation within a group is a draw from a mixture model and where it is desirable to share mixture components between groups, and considers a hierarchical model, specifically one in which the base measure for the childDirichlet processes is itself distributed according to a Dirichlet process.
Empirical Methods in Natural Language ProcessingTextRank: Bringing Order into Text
3,347 Citations2004Rada Mihalcea, Paul Tarau
TextRank, a graph-based ranking model for text processing, is introduced and it is shown how this model can be successfully used in natural language applications.
IBM Journal of Research and DevelopmentThe Automatic Creation of Literature Abstracts
3,211 Citations1958H. P. Luhn
In the exploratory research described, the complete text of an article in machine-readable form is scanned by an IBM 704 data-processing machine and analyzed in accordance with a standard program.
IEEE Signal Processing MagazineThe expectation-maximization algorithm
3,138 Citations1996Todd K. Moon
The EM (expectation-maximization) algorithm is ideally suited to problems of parameter estimation, in that it produces maximum-likelihood (ML) estimates of parameters when there is a many-to-one mapping from an underlying distribution to the distribution governing the observation.
Computational LinguisticsA maximum entropy approach to natural language processing
3,120 Citations1996Adam Berger, Vincent J. Della Pietra +1 more
A maximum-likelihood approach for automatically constructing maximum entropy models is presented and how to implement this approach efficiently is described, using as examples several problems in natural language processing.
Machine LearningOn the Optimality of the Simple Bayesian Classifier under Zero-One Loss
3,072 Citations1997Pedro Domingos, Michael J. Pazzani
The Bayesian classifier is shown to be optimal for learning conjunctions and disjunctions, even though they violate the independence assumption, and will often outperform more powerful classifiers for common training set sizes and numbers of attributes, even if its bias is a priori much less appropriate to the domain.
Distant supervision for relation extraction without labeled data
2,913 Citations2009Mike D. Mintz, Steven Bills +2 more
This work investigates an alternative paradigm that does not require labeled corpora, avoiding the domain dependence of ACE-style algorithms, and allowing the use of corpora of any size.
Transductive Inference for Text Classification using Support Vector Machines
2,717 Citations1999Thorsten Joachims
An analysis of why TSVMs are well suited for text classi(cid:12)cation is presented, and an algorithm for training TSVMs e(cid:14)-ciently, handling 10,000 examples and more is proposed.
Accurate methods for the statistics of surprise and coincidence
2,688 Citations1993Ted Dunning
Machine LearningMarkov logic networks
2,678 Citations2006Matthew Richardson, Pedro Domingos
Experiments with a real-world database and knowledge base in a university domain illustrate the promise of this approach to combining first-order logic and probabilistic graphical models in a single representation.
A gentle tutorial of the em algorithm and its application to parameter estimation for Gaussian mixture and hidden Markov models
2,508 Citations1998Jeffrey A. Bilmes
Dynamic topic models
2,322 Citations2006David M. Blei, John Lafferty
A family of probabilistic time series models is developed to analyze the time evolution of topics in large document collections, and dynamic topic models provide a qualitative window into the contents of a large document collection.
Journal of Computational and Graphical StatisticsMarkov Chain Sampling Methods for Dirichlet Process Mixture Models
2,212 Citations2000Radford M. Neal
arXiv (Cornell University)Probabilistic Latent Semantic Analysis
2,092 Citations2013Thomas Hofmann
This work proposes a widely applicable generalization of maximum likelihood model fitting by tempered EM, based on a mixture decomposition derived from a latent class model which results in a more principled approach which has a solid foundation in statistics.
Lecture notes in computer scienceNaive (Bayes) at forty: The independence assumption in information retrieval
2,092 Citations1998David Lewis
The naive Bayes classifier, currently experiencing a renaissance in machine learning, has long been a core technique in information retrieval, and some of the variations used for text retrieval and classification are reviewed.
Journal of the American Society for Information ScienceRelevance weighting of search terms
2,068 Citations1976Stephen Robertson, Karen Spärck Jones
This paper examines statistical techniques for exploiting relevance information to weight search terms using information about the distribution of index terms in documents in general and shows that specific weighted search methods are implied by a general probabilistic theory of retrieval.
Society for Industrial and Applied Mathematics eBooksIterative Methods for Optimization
2,030 Citations1999C. T. Kelley
Iterative Methods for Optimization does more than cover traditional gradient-based optimization: it is the first book to treat sampling methods, including the Hooke& Jeeves, implicit filtering, MDS, and Nelder& Mead schemes in a unified way.
The MIT Press eBooksAnalysis of Representations for Domain Adaptation
2,023 Citations2007Shai Ben-David, John Blitzer +2 more
The theory illustrates the tradeoffs inherent in designing a representation for domain adaptation and gives a new justification for a recently proposed model which explicitly minimizes the difference between the source and target domains, while at the same time maximizing the margin of the training set.
Elsevier eBooksNewsWeeder: Learning to Filter Netnews
2,022 Citations1995Ken Lang
The results show that a learning algorithm based on the Minimum Description Length (MDL) principle was able to raise the percentage of interesting articles to be shown to users from 14% to 52% on average.
The foundations of cost-sensitive learning
1,835 Citations2001Charles Elkan
It is argued that changing the balance of negative and positive training examples has little effect on the classifiers produced by standard Bayesian and decision tree learning methods, and the recommended way of applying one of these methods is to learn a classifier from the training set and then to compute optimal decisions explicitly using the probability estimates given by the classifier.
Synthesis lectures on artificial intelligence and machine learningIntroduction to Semi-Supervised Learning
1,804 Citations2009Xiaojin Zhu, Andrew B. Goldberg
This introductory book presents some popular semi-supervised learning models, including self-training, mixture models, co-training and multiview learning, graph-based methods, and semi- supervised support vector machines, and discusses their basic mathematical formulation.
Society for Industrial and Applied Mathematics eBooksNumerical Methods for Large Eigenvalue Problems
1,656 Citations2011Yousef Saad
Future Generation Computer SystemsMining generalized association rules
1,617 Citations1997Ramakrishnan Srikant, Rakesh Agrawal
A new interest-measure for rules which uses the information in the taxonomy is presented, and given a user-specified “minimum-interest-level”, this measure prunes a large number of redundant rules.
Transformation-based error-driven learning and natural language processing: a case study in part-of-speech tagging
1,531 Citations1995Eric Brill
This paper describes a simple rule-based approach to automated learning of linguistic knowledge that has been shown for a number of tasks to capture information in a clearer and more direct fashion without a compromise in performance.
PubMedMixed Membership Stochastic Blockmodels.
1,526 Citations2008Edoardo M. Airoldi, David M. Blei +2 more
This paper describes a latent variable model of such data called the mixed membership stochastic blockmodel, which extends blockmodels for relational data to ones which capture mixed membership latent relational structure, thus providing an object-specific low-dimensional representation.
SIAM ReviewUsing Linear Algebra for Intelligent Information Retrieval
1,484 Citations1995Michael W. Berry, Susan Dumais +1 more
A lexical match between words in users’ requests and those in or assigned to documents in a database helps retrieve textual materials from scientific databases.
Journal of the ACMNew Methods in Automatic Extracting
1,469 Citations1969H. P. Edmundson
New methods of automatically extracting documents for screening purposes, i.e. the computer selection of sentences having the greatest potential for conveying to the reader the substance of the document, indicate that the three newly proposed components dominate the frequency component in the production of better extracts.
arXiv (Cornell University)Loopy Belief Propagation for Approximate Inference: An Empirical Study
1,468 Citations2013Kevin P. Murphy, Yair Weiss +1 more
Inductive learning algorithms and representations for text categorization
1,465 Citations1998Susan Dumais, John Platt +2 more
A comparison of the effectiveness of five different automatic learning algorithms for text categorization in terms of learning speed, realtime classification speed, and classification accuracy is compared.
IEEE Transactions on Neural NetworksSupport vector machines for spam categorization
1,457 Citations1999Harris Drucker, Donghui Wu +1 more
The use of support vector machines in classifying e-mail as spam or nonspam is studied by comparing it to three other classification algorithms: Ripper, Rocchio, and boosting decision trees, which found SVM's performed best when using binary features.
Bayesian AnalysisVariational inference for Dirichlet process mixtures
1,446 Citations2006David M. Blei, Michael I. Jordan
A variational inference algorithm forDP mixtures is presented and experiments that compare the algorithm to Gibbs sampling algorithms for DP mixtures of Gaussians and present an application to a large-scale image analysis problem are presented.
ArXiv.orgFrustratingly Easy Domain Adaptation
1,395 Citations2009Hal Daumé
This work describes an approach to domain adaptation that is appropriate exactly in the case when one has enough “target” data to do slightly better than just using only “source’ data.
Machine LearningLearning Quickly When Irrelevant Attributes Abound: A New Linear-Threshold Algorithm
1,376 Citations1988Nick Littlestone
This work presents one such algorithm that learns disjunctive Boolean functions, along with variants for learning other classes of Boolean functions.
Journal of Machine Learning ResearchA Framework for Learning Predictive Structures from Multiple Tasks and Unlabeled Data
1,372 Citations2005Rie Kubota Ando, Tong Zhang
This paper presents a general framework in which the structural learning problem can be formulated and analyzed theoretically, and relate it to learning with unlabeled data, and algorithms for structural learning will be proposed, and computational issues will be investigated.
Message Understanding Conference-6
1,365 Citations1996Ralph Grishman, Beth Sundheim
MUC-6 introduced several innovations over prior MUCs, most notably in the range of different tasks for which evaluations were conducted and the motivations for the new format.
Maximum Entropy Markov Models for Information Extraction and Segmentation
1,333 Citations2000Andrew McCallum, Dayne Freitag +1 more
A new Markovian sequence model is presented that allows observations to be represented as arbitrary overlapping features (such as word, capitalization, formatting, part-of-speech), and defines the conditional probability of state sequences given observation sequences.
Open information extraction from the web
1,323 Citations2007Michele Banko, Michael Cafarella +3 more
arXiv (Cornell University)Supervised Topic Models
1,315 Citations2010David M. Blei, Jon McAuliffe
The supervised latent Dirichlet allocation (sLDA) model, a statistical model of labelled documents, is introduced, which derives a maximum-likelihood procedure for parameter estimation, which relies on variational approximations to handle intractable posterior expectations.
Uncertainty in Artificial IntelligenceThe author-topic model for authors and documents
1,313 Citations2004Michal Rosen‐Zvi, Thomas Griffiths +2 more
Maximum Entropy Model for Part-Of-Speech Tagging
1,276 Citations1996Adwait Ratnaparkhi
A statistical model which trains from a corpus annotated with Part Of Speech tags and assigns them to previously unseen text with state of the art accuracy and discusses the corpus consistency problems discovered during the implementation of these features.
A Probabilistic Analysis of the Rocchio Algorithm with TFIDF for Text Categorization
1,264 Citations1997Thorsten Joachims
A Probabilistic analysis of the Rocchio relevance feedback algorithm, one of the most popular learning methods from information retrieval, is presented in a text categorization framework and suggests that the probabilistic algorithms are preferable to the heuristic Rocchio classifier.
Text, speech and language technologyText Chunking Using Transformation-Based Learning
1,257 Citations1999Lance Ramshaw, Mitchell P. Marcus
This work has shown that the transformation-based learning approach can be applied at a higher level of textual interpretation for locating chunks in the tagged text, including non-recursive “baseNP” chunks.
Relational learning via collective matrix factorization
1,235 Citations2008Ajit Singh, Geoffrey J. Gordon
This model generalizes several existing matrix factorization methods, and therefore yields new large-scale optimization algorithms for these problems, which can handle any pairwise relational schema and a wide variety of error models.
<i>Snowball</i>
1,169 Citations2000Eugene Agichtein, Luis Gravano
This paper develops a scalable evaluation methodology and metrics for the task, and presents a thorough experimental evaluation of Snowball and comparable techniques over a collection of more than 300,000 newspaper documents.
Early results for named entity recognition with conditional random fields, feature induction and web-enhanced lexicons
1,161 Citations2003Andrew McCallum, Wei Li
This work has shown that conditionally-trained models, such as conditional maximum entropy models, handle inter-dependent features of greedy sequence modeling in NLP well.
Identifying Relations for Open Information Extraction
1,152 Citations2011Anthony Fader, Stephen Soderland +1 more
Two simple syntactic and lexical constraints on binary relations expressed by verbs are introduced in the ReVerb Open IE system, which more than doubles the area under the precision-recall curve relative to previous extractors such as TextRunner and woepos.
OPAL (Open@LaTrobe) (La Trobe University)A correlated topic model of Science
1,137 Citations2018David M. Blei, John Lafferty
LDA-based document models for ad-hoc retrieval
1,086 Citations2006Xing Wei, W. Bruce Croft
This paper proposes an LDA-based document model within the language modeling framework, and evaluates it on several TREC collections, and shows that improvements over retrieval using cluster-based models can be obtained with reasonable efficiency.
Topic modeling
1,057 Citations2006Hanna Wallach
A hierarchical generative probabilistic model that incorporates both n-gram statistics and latent topic variables by extending a unigram topic model to include properties of a hierarchical Dirichlet bigram language model is explored.
Wrapper induction for information extraction
1,044 Citations1997Nicholas Kushmerick, Daniel S. Weld
This work introduces wrapper induction, a method for automatically constructing wrappers, and identifies hlrt, a wrapper class that is e(cid:14)ciently learnable, yet expressive enough to handle 48% of a recently surveyed sample of Internet resources.
Lecture notes in computer scienceExtracting Patterns and Relations from the World Wide Web
1,007 Citations1999Sergey Brin
This paper presents a technique which exploits the duality between sets of patterns and relations to grow the target relation starting from a small sample and uses it to extract a relation of (author,title) pairs from the World Wide Web.
On the Equivalence of Nonnegative Matrix Factorization and Spectral Clustering
977 Citations2005Chris Ding, Xiaofeng He +1 more
This work shows that (1) W = HH T is equivalent to Kernel K -means clustering and the Laplacian-based spectral clustering and (2) X = FG T is equivalent to simultaneous clustering of rows and columns of a bipartite graph.
Neural Information Processing SystemsHierarchical Topic Models and the Nested Chinese Restaurant Process
938 Citations2003Thomas L. Griffiths, Michael I. Jordan +2 more
A Bayesian approach is taken to generate an appropriate prior via a distribution on partitions that allows arbitrarily large branching factors and readily accommodates growing data collections.
Generic text summarization using relevance measure and latent semantic analysis
847 Citations2001Yihong Gong, Xin Liu
This paper proposes two generic text summarization methods that create text summaries by ranking and extracting sentences from the original documents, and uses the latent semantic analysis technique to identify semantically important sentences, for summary creations.
Natural language processingAutomatic Summarization
839 Citations2001Inderjeet Mani
The challenges that remain open, in particular the need for language generation and deeper semantic understanding of language that would be necessary for future advances in the field are discussed.
Dependency tree kernels for relation extraction
825 Citations2004Aron Culotta, Jeffrey Sorensen
This work extends previous work on tree kernels to estimate the similarity between the dependency trees of sentences, and uses this kernel within a Support Vector Machine to detect and classify relations between entities in the Automatic Content Extraction (ACE) corpus of news articles.
Natural Language EngineeringTechnical terminology: some linguistic properties and an algorithm for identification in text
817 Citations1995John S. Justeson, Slava M. Katz
The paper presents a terminology indentification algorithm that recovers a high proportion of the technical terms in a text, and a high proportaion of the recovered strings are vaild technical terms.
Semi-supervised Clustering by Seeding
804 Citations2002Sugato Basu, Arindam Banerjee +1 more
Modeling online reviews with multi-grain topic models
792 Citations2008Ivan Titov, Ryan McDonald
This paper presents a novel framework for extracting ratable aspects of objects from online user reviews and argues that multi-grain models are more appropriate for this task since standard models tend to produce topics that correspond to global properties of objects rather than aspects of an object that tend to be rated by a user.
Cross-domain sentiment classification via spectral feature alignment
789 Citations2010Sinno Jialin Pan, Xiaochuan Ni +3 more
This work develops a general solution to sentiment classification when the authors do not have any labels in a target domain but have some labeled data in a different domain, regarded as source domain and proposes a spectral feature alignment (SFA) algorithm to align domain-specific words from different domains into unified clusters, with the help of domain-independent words as a bridge.
Machine LearningAn Algorithm that Learns What's in a Name
784 Citations1999Daniel M. Bikel, Richard Schwartz +1 more
IdentiFinderTM, a hidden Markov model that learns to recognize and classify names, dates, times, and numerical quantities, is evaluated and is competitive with approaches based on handcrafted rules on mixed case text and superior on text where case information is not available.
Singapore Management University Institutional Knowledge (InK) (Singapore Management University)Instance Weighting for Domain Adaptation in NLP
754 Citations2007Jing Jiang, ChengXiang Zhai
This paper formally analyze and characterize the domain adaptation problem from a distributional view, and shows that there are two distinct needs for adaptation, corresponding to the different distributions of instances and classification functions in the source and the target domains.
Cross-domain video concept detection using adaptive svms
697 Citations2007Jun Yang, Rong Yan +1 more
This paper proposes Adaptive Support Vector Machines (A-SVMs) as a general method to adapt one or more existing classifiers of any type to the new dataset and outperforms several baseline and competing methods in terms of classification accuracy and efficiency in cross-domain concept detection in the TRECVID corpus.
Journal of the ACMThe nested chinese restaurant process and bayesian nonparametric inference of topic hierarchies
667 Citations2010David M. Blei, Thomas L. Griffiths +1 more
An application to information retrieval in which documents are modeled as paths down a random tree, and the preferential attachment dynamics of the nCRP leads to clustering of documents according to sharing of topics at multiple levels of abstraction.
A comparison of algorithms for maximum entropy parameter estimation
667 Citations2002Robert Malouf
A number of algorithms for estimating the parameters of ME models are considered, including iterative scaling, gradient ascent, conjugate gradient, and variable metric methods.
ACM SIGIR ForumOn-Line New Event Detection and Tracking
653 Citations2017James Allan, Ron Papka +1 more
Streaming-data algorithms for high-quality clustering
579 Citations2003Liadan O'Callaghan, Nita Mishra +3 more
This work describes a streaming algorithm that effectively clusters large data streams and provides empirical evidence of the algorithm's performance on synthetic and real data streams.
ScholarWorks@UMassAmherst (University of Massachusetts Amherst)Rethinking LDA: Why Priors Matter
563 Citations2009Hanna Wallach, David Mimno +1 more
The prior structure advocated substantially increases the robustness of topic models to variations in the number of topics and to the highly skewed word frequency distributions common in natural language.
…
