An Introduction to Conditional Random Fields
Foundations and Trends® in Machine LearningPublished 1 January 2012Open access
Charles Sutton
Citations529
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
Abstract
An Introduction to Conditional Random Fields provides a comprehensive tutorial aimed at application-oriented practitioners seeking to apply CRFs. The monograph does not assume previous knowledge of graphical modeling, and so is intended to be useful to practitioners in a wide variety of fields.
Keywords
Computer Science
Proceedings of the IEEEGradient-based learning applied to document recognition
58,219 Citations1998Yann LeCun, Léon Bottou +2 more
This paper reviews various methods applied to handwritten character recognition and compares them on a standard handwritten digit recognition task, and Convolutional neural networks are shown to outperform all other techniques.
Journal of Machine Learning ResearchLatent dirichlet allocation
27,049 Citations2003David M. Blei, Andrew Y. Ng +1 more
Proceedings of the IEEEA tutorial on hidden Markov models and selected applications in speech recognition
22,785 Citations1989L. R. Rabiner
IEEE Transactions on Pattern Analysis and Machine IntelligenceStochastic Relaxation, Gibbs Distributions, and the Bayesian Restoration of Images
17,980 Citations1984Stuart Geman, Donald Geman
The analogy between images and statistical mechanics systems is made and the analogous operation under the posterior distribution yields the maximum a posteriori (MAP) estimate of the image given the degraded observations, creating a highly parallel ``relaxation'' algorithm for MAP estimation.
Probabilistic Reasoning in Intelligent Systems: Networks of Plausible Inference
16,927 Citations1988Judea Pearl
The author provides a coherent explication of probability as a language for reasoning with partial belief and offers a unifying perspective on other AI approaches to uncertainty, such as the Dempster-Shafer formalism, truth maintenance systems, and nonmonotonic logic.
Object recognition from local scale-invariant features
16,195 Citations1999David Lowe
Experimental results show that robust object recognition can be achieved in cluttered partially occluded images with a computation time of under 2 seconds.
Cambridge University Press eBooksNumerical optimization
14,107 Citations2021W. John Braun, Duncan J. Murdoch
ScholarlyCommons (University of Pennsylvania)Conditional Random Fields: Probabilistic Models for Segmenting and Labeling Sequence Data
12,978 Citations2001John Lafferty, Andrew McCallum +1 more
This work presents iterative parameter estimation algorithms for conditional random fields and compares the performance of the resulting models to HMMs and MEMMs on synthetic and natural-language data.
IEEE Transactions on Pattern Analysis and Machine IntelligenceObject Detection with Discriminatively Trained Part-Based Models
9,976 Citations2009Pedro F. Felzenszwalb, Ross Girshick +2 more
An object detection system based on mixtures of multiscale deformable part models that is able to represent highly variable object classes and achieves state-of-the-art results in the PASCAL object detection challenges is described.
The Annals of Mathematical StatisticsA Stochastic Approximation Method
9,581 Citations1951Herbert Robbins, Sutton Monro
Building a Large Annotated Corpus of English: The Penn Treebank
7,528 Citations1993Mitchell P. Marcus
Journal of the American Statistical AssociationSampling-Based Approaches to Calculating Marginal Densities
6,616 Citations1990Alan E. Gelfand, A. F. M. Smith
Stochastic substitution, the Gibbs sampler, and the sampling-importance-resampling algorithm can be viewed as three alternative sampling- (or Monte Carlo-) based approaches to the calculation of numerical estimates of marginal probability distributions.
IEEE Transactions on Information TheoryFactor graphs and the sum-product algorithm
6,481 Citations2001Frank R. Kschischang, Brendan J. Frey +1 more
A generic message-passing algorithm, the sum-product algorithm, that operates in a factor graph, that computes-either exactly or approximately-various marginal functions derived from the global function.
Probabilistic graphical models : principles and techniques
6,434 Citations2009Daniel L. Koller, Nir Friedman
The framework of probabilistic graphical models, presented in this book, provides a general approach for causal reasoning and decision making under uncertainty, allowing interpretable models to be constructed and then manipulated by reasoning algorithms.
Journal of the Royal Statistical Society Series B (Statistical Methodology)Spatial Interaction and the Statistical Analysis of Lattice Systems
6,179 Citations1974Julian Besag
TechnometricsMonte Carlo Statistical Methods
5,612 Citations2000Hoon Kim, Christian P. Robert +1 more
Statistics and ComputingWinBUGS - A Bayesian modelling framework: Concepts, structure, and extensibility
5,395 Citations2000David J. Lunn, Andrew C. Thomas +2 more
How and why various modern computing concepts, such as object-orientation and run-time linking, feature in the software's design are discussed and how the framework may be extended.
Neural ComputationTraining Products of Experts by Minimizing Contrastive Divergence
5,017 Citations2002Geoffrey E. Hinton
A product of experts (PoE) is an interesting candidate for a perceptual system in which rapid inference is vital and generation is unnecessary because it is hard even to approximate the derivatives of the renormalization term in the combination rule.
Interactive graph cuts for optimal boundary & region segmentation of objects in N-D images
3,676 Citations2002Yuri Boykov, Marie‐Pierre Jolly
A new technique for general purpose interactive segmentation of N-dimensional images where the user marks certain pixels as "object" or "background" to provide hard constraints for segmentation.
now publishers, Inc. eBooksGraphical Models, Exponential Families, and Variational Inference
3,154 Citations2007Martin J. Wainwright, Michael I. Jordan
Incorporating non-local information into information extraction systems by Gibbs sampling
3,035 Citations2005Jenny Rose Finkel, Trond Grenager +1 more
By using simulated annealing in place of Viterbi decoding in sequence models such as HMMs, CMMs, and CRFs, it is possible to incorporate non-local structure while preserving tractable inference.
Feature-rich part-of-speech tagging with a cyclic dependency network
2,851 Citations2003Kristina Toutanova, Dan Klein +2 more
A new part-of-speech tagger is presented that demonstrates the following ideas: explicit use of both preceding and following tag contexts via a dependency network representation, broad use of lexical features, and effective use of priors in conditional loglinear models.
Machine LearningMarkov logic networks
2,678 Citations2006Matthew Richardson, Pedro Domingos
Experiments with a real-world database and knowledge base in a university domain illustrate the promise of this approach to combining first-order logic and probabilistic graphical models in a single representation.
The Annals of Mathematical StatisticsStochastic Estimation of the Maximum of a Regression Function
2,143 Citations1952J. Kiefer, J. Wolfowitz
Discriminative training methods for hidden Markov models
1,889 Citations2002Michael Collins
Experimental results on part-of-speech tagging and base noun phrase chunking are given, in both cases showing improvements over results for a maximum-entropy tagger.
On Discriminative vs. Generative Classifiers: A comparison of logistic regression and naive Bayes
1,887 Citations2001Andrew Y. Ng, Michael I. Jordan
It is shown, contrary to a widely-held belief that discriminative classifiers are almost always to be preferred, that there can often be two distinct regimes of performance as the training set size is increased, one in which each algorithm does better.
IEEE Transactions on Information TheoryConstructing Free-Energy Approximations and Generalized Belief Propagation Algorithms
1,627 Citations2005Jonathan S. Yedidia, William T. Freeman +1 more
This work explains how to obtain region-based free energy approximations that improve the Bethe approximation, and corresponding generalized belief propagation (GBP) algorithms, and describes empirical results showing that GBP can significantly outperform BP.
Journal of the Royal Statistical Society Series D (The Statistician)Statistical Analysis of Non-Lattice Data
1,575 Citations1975Julian Besag
Feature selection, <i>L</i><sub>1</sub> vs. <i>L</i><sub>2</sub> regularization, and rotational invariance
1,555 Citations2004Andrew Y. Ng
A lower-bound is given showing that any rotationally invariant algorithm---including logistic regression with L1 regularization, SVMs, and neural networks trained by backpropagation---has a worst case sample complexity that grows at least linearly in the number of irrelevant features.
Computer science workbenchMarkov Random Field Modeling in Image Analysis
1,542 Citations2001Stan Z. Li
This detailed and thoroughly enhanced third edition presents a comprehensive study / reference to theories, methodologies and recent developments in solving computer vision problems based on MRFs, statistics and optimisation.
Mathematical ProgrammingPegasos: primal estimated sub-gradient solver for SVM
1,524 Citations2010Shai Shalev‐Shwartz, Yoram Singer +2 more
Online Passive-Aggressive Algorithms
1,449 Citations2006Koby Crammer, Ofer Dekel +3 more
Maximum Entropy Markov Models for Information Extraction and Segmentation
1,333 Citations2000Andrew McCallum, Dayne Freitag +1 more
A new Markovian sequence model is presented that allows observations to be represented as arbitrary overlapping features (such as word, capitalization, formatting, part-of-speech), and defines the conditional probability of state sequences given observation sequences.
Support vector machine learning for interdependent and structured output spaces
1,260 Citations2004Ioannis Tsochantaridis, Thomas Hofmann +2 more
This paper proposes to generalize multiclass Support Vector Machine learning in a formulation that involves features extracted jointly from inputs and outputs, and demonstrates the versatility and effectiveness of the method on problems ranging from supervised grammar learning and named-entity recognition, to taxonomic text classification and sequence alignment.
The MIT Press eBooksMap-Reduce for Machine Learning on Multicore
1,253 Citations2007Cheng-Tao Chu, Sang Kyun Kim +5 more
This work shows that algorithms that fit the Statistical Query model can be written in a certain "summation form," which allows them to be easily parallelized on multicore computers and shows basically linear speedup with an increasing number of processors.
Shallow parsing with conditional random fields
1,251 Citations2003Fei Sha, Fernando Pereira
This work shows how to train a conditional random field to achieve performance as good as any reported base noun-phrase chunking method on the CoNLL task, and better than any reported single model.
Max-Margin Markov Networks
1,251 Citations2003Ben Taskar, Carlos Guestrin +1 more
Maximum margin Markov (M3) networks incorporate both kernels, which efficiently deal with high-dimensional features, and the ability to capture correlations in structured data, and a new theoretical bound for generalization in structured domains is provided.
Early results for named entity recognition with conditional random fields, feature induction and web-enhanced lexicons
1,161 Citations2003Andrew McCallum, Wei Li
This work has shown that conditionally-trained models, such as conditional maximum entropy models, handle inter-dependent features of greedy sequence modeling in NLP well.
Lecture notes in computer scienceTextonBoost: Joint Appearance, Shape and Context Modeling for Multi-class Object Recognition and Segmentation
1,153 Citations2006Jamie Shotton, John Winn +2 more
A new approach to learning a discriminative model of object classes, incorporating appearance, shape and context information efficiently, is proposed, which is used for automatic visual recognition and semantic segmentation of photographs.
Multitask learning
1,118 Citations1998Rich Caruana
Semi-supervised Learning by Entropy Minimization.
1,021 Citations2005Yves Grandvalet, Yoshua Bengio
This framework, which motivates minimum entropy regularization, enables to incorporate unlabeled data in the standard supervised learning, and includes other approaches to the semi-supervised problem as particular or limiting cases.
Pegasos
982 Citations2007Shai Shalev‐Shwartz, Yoram Singer +1 more
A simple and effective stochastic sub-gradient descent algorithm for solving the optimization problem cast by Support Vector Machines, which is particularly well suited for large text classification problems, and demonstrates an order-of-magnitude speedup over previous SVM learning methods.
IEEE Journal on Selected Areas in CommunicationsTurbo decoding as an instance of Pearl's "belief propagation" algorithm
905 Citations1998Robert J. McEliece, David Mackay +1 more
It is shown that Pearl's algorithm can be used to routinely derive previously known iterative, but suboptimal, decoding algorithms for a number of other error-control systems, including Gallager's low-density parity-check codes, serially concatenated codes, and product codes.
International Journal of Computer VisionRobust Higher Order Potentials for Enforcing Label Consistency
893 Citations2009Pushmeet Kohli, Ľubor Ladický +1 more
This paper proposes a novel framework for labelling problems which is able to combine multiple segmentations in a principled manner based on higher order conditional random fields and uses potentials defined on sets of pixels generated using unsupervised segmentation algorithms.
Multiscale conditional random fields for image labeling
843 Citations2004Xuming He, Richard S. Zemel +1 more
An approach to include contextual features for labeling images, in which each pixel is assigned to one of a finite set of labels, are incorporated into a probabilistic framework, which combines the outputs of several components.
Maximum mutual information estimation of hidden Markov model parameters for speech recognition
821 Citations2005L.R. Bahl, Peter Brown +2 more
A method for estimating the parameters of hidden Markov models of speech is described and recognition results are presented comparing this method with maximum likelihood estimation.
A Tutorial on Energy-Based Learning
793 Citations2006Yann LeCun, Sumit Chopra +3 more
The EBM approach provides a common theoretical framework for many learning models, including traditional discr iminative and generative approaches, as well as graph-transformer networks, co nditional random fields, maximum margin Markov networks, and several manifold learning methods.
IEEE Transactions on Information TheoryThe generalized distributive law
754 Citations2000S.M. Aji, Robert J. McEliece
Although this algorithm is guaranteed to give exact answers only in certain cases (the "junction tree" condition), unfortunately not including the cases of GTW with cycles or turbo decoding, there is much experimental evidence, and a few theorems, suggesting that it often works approximately even when it is not supposed to.
Dynamic conditional random fields
746 Citations2004Charles Sutton, Khashayar Rohanimanesh +1 more
Mathematical ProgrammingRepresentations of quasi-Newton matrices and their use in limited memory methods
746 Citations1994Richard H. Byrd, Jorge Nocedal +1 more
This work derives compact representations of BFGS and symmetric rank-one matrices for optimization and presents a compact representation of the matrices generated by Broyden's update for solving systems of nonlinear equations.
Applying Conditional Random Fields to Japanese Morphological Analysis
723 Citations2004Taku Kudo, Kaoru Yamamoto +1 more
This paper shows how CRFs can be applied to situations where word boundary ambiguity exists, and confirms that CRFs offer a solution to the long-standing problems in corpus-based or statistical Japanese morphological analysis.
A comparison of algorithms for maximum entropy parameter estimation
667 Citations2002Robert Malouf
A number of algorithms for estimating the parameters of ME models are considered, including iterative scaling, gradient ascent, conjugate gradient, and variable metric methods.
Discriminative probabilistic models for relational data
637 Citations2002Ben Taskar, Pieter Abbeel +1 more
Maximum margin planning
630 Citations2006Nathan Ratliff, J. Andrew Bagnell +1 more
This work learns mappings from features to cost so an optimal policy in an MDP with these cost mimics the expert's behavior, and demonstrates a simple, provably efficient approach to structured maximum margin learning, based on the subgradient method, that leverages existing fast algorithms for inference.
Learning structural SVMs with latent variables
627 Citations2009Chun-Nam Yu, Thorsten Joachims
A large-margin formulation and algorithm for structured output prediction that allows the use of latent variables and the generality and performance of the approach is demonstrated through three applications including motiffinding, noun-phrase coreference resolution, and optimizing precision at k in information retrieval.
Semi-Markov Conditional Random Fields for Information Extraction
617 Citations2004Sunita Sarawagi, William W. Cohen
Intuitively, a semi-CRF on an input sequence x outputs a "segmentation" of x, in which labels are assigned to segments rather than to individual elements of xi, and transitions within a segment can be non-Markovian.
Introduction to the bio-entity recognition task at JNLPBA
579 Citations2004Jin-Dong Kim, Tomoko Ohta +3 more
The JNLPBA shared task of bio-entity recognition using an extended version of the GENIA version 3 named entity corpus of MEDLINE abstracts is described and a general discussion of the approaches taken by participating systems is presented.
BMC BioinformaticsOverview of BioCreAtIvE: critical assessment of information extraction for biology
542 Citations2005Lynette Hirschman, Alexander Yeh +2 more
The first BioCreAtIvE assessment provided state-of-the-art performance results for a basic task (gene name finding and normalization), where the best systems achieved a balanced 80% precision / recall or better, which potentially makes them suitable for real applications in biology.
Scalable training of<i>L</i><sup>1</sup>-regularized log-linear models
541 Citations2007Galen Andrew, Jianfeng Gao
This work presents an algorithm Orthant-Wise Limited-memory Quasi-Newton (OWL-QN), based on L-BFGS, that can efficiently optimize the L1-regularized log-likelihood of log-linear models with millions of parameters.
The MIT Press eBooksPredicting Structured Data
533 Citations2007
This volume presents the state of the art in machine learning algorithms and theory in this novel field and discusses applications as diverse as machine translation, document markup, computational biology, and information extraction, providing a timely overview of an exciting field.
IEEE Transactions on Pattern Analysis and Machine IntelligenceHidden Conditional Random Fields
490 Citations2007Ariadna Quattoni, Sybor Wang +3 more
A discriminative latent variable model for classification problems in structured domains where inputs can be represented by a graph of local observations and a hidden-state conditional random field framework learns a set of latent variables conditioned on local features.
Hidden Markov support vector machines
469 Citations2003Yasemin Altün, Ioannis Tsochantaridis +1 more
This paper presents a novel discriminative learning technique for label sequences based on a combination of the two most successful learning algorithms, Support Vector Machines and Hidden Markov Models which it is called HM-SVMs and handles dependencies between neighboring labels using Viterbi decoding.
Chinese segmentation and new word detection using conditional random fields
468 Citations2004Fuchun Peng, Fangfang Feng +1 more
The ability of linear-chain conditional random fields (CRFs) to perform robust and accurate Chinese word segmentation by providing a principled framework that easily supports the integration of domain knowledge in the form of multiple lexicons of characters and words is demonstrated.
Computer applications in the biosciencesABNER: an open source tool for automatically tagging genes, proteins and other entity names in text
467 Citations2005Burr Settles
ScholarlyCommons (University of Pennsylvania)Posterior Regularization for Structured Latent Variable Models
462 Citations2010Kuzman Ganchev, Joäo Graça +2 more
Divergence measures and message passing
460 Citations2005Thomas P. Minka
This paper presents a unifying view of messagepassing algorithms, as methods to approximate a complex Bayesian network by a simpler network with minimum information divergence.
Composite Likelihood Methods
458 Citations2012Harry Joe, Nancy Reid +2 more
This issue includes two long overview papers, one of which is devoted to applications in statistical genetics; several papers developing new theory for inference based on composite likelihood; new results in the application of composite likelihood to time series, spatial processes, longitudinal data and missing data.
Machine LearningSearch-based structured prediction
436 Citations2009Hal Daumé, John Langford +1 more
Searn is an algorithm for integrating search and learning to solve complex structured prediction problems such as those that occur in natural language, speech, computational biology, and vision and comes with a strong, natural theoretical guarantee: good performance on the derived classification problems implies goodperformance on the structured prediction problem.
Collective multi-label classification
416 Citations2005Nadia Ghamrawi, Andrew McCallum
This paper explores multi-label conditional random field (CRF) classification models that directly parameterize label co-occurrences in multi-label classification.
International Journal of Computer VisionDiscriminative Random Fields
392 Citations2006Sanjiv Kumar, Martial Hebert
This work presents Discriminative Random Fields (DRFs) to model spatial interactions in images in a discriminative framework based on the concept of Conditional Random Fields proposed by lafferty et al.(2001).
Identifying sources of opinions with conditional random fields and extraction patterns
366 Citations2005Yejin Choi, Claire Cardie +2 more
This work adopts a hybrid approach that combines Conditional Random Fields (Lafferty et al., 2001) and a variation of AutoSlog (Riloff, 1996a), and shows that the combination of these two methods performs better than either one alone.
arXiv (Cornell University)Efficiently Inducing Features of Conditional Random Fields
366 Citations2012Andrew McCallum
Table extraction using conditional random fields
355 Citations2003David Pinto, Andrew McCallum +2 more
Unlike HMMs, CRFs support the use of many rich and overlapping layout and language features, and as a result, they perform significantly better, and are compared with hidden Markov models (HMMs).
Conditional Random Fields for Object Recognition
338 Citations2004Ariadna Quattoni, Michael Collins +1 more
An extension of the CRF framework that incorporates hidden variables and combines class conditional CRFs into a unified framework for part-based object recognition is proposed, which allows the assumption of conditional independence of the observed data to be relaxed.
Applied Physics Letters10.1162/jmlr.2003.3.4-5.951
323 Citations2000
This paper describes a family of additive ultraconservative algorithms where each algorithm in the family updates its prototypes by finding a feasible solution for a set of linear constraints that depend on the instantaneous similarity-scores.
Hidden conditional random fields for phone classification
302 Citations2005Asela Gunawardana, Milind Mahajan +2 more
This paper presents the results on the TIMIT phone classification task and shows that HCRFs outperforms comparable ML and CML/MMI trained HMMs and has the ability to handle complex features without any change in training procedure.
Practical Very Large Scale CRFs
299 Citations2010Thomas Lavergne, Olivier Cappé +1 more
This paper addresses the issue of training very large CRFs, containing up to hundreds output labels and several billion features, and indicates that efficiency stems here from the sparsity induced by the use of a l penalty term.
Edinburgh Research Explorer (University of Edinburgh)Dynamic Conditional Random Fields: Factorized Probabilistic Models for Labeling and Segmenting Sequence Data
292 Citations2007Charles Sutton, Andrew McCallum +1 more
On a natural-language chunking task, it is shown that a DCRF performs better than a series of linear-chain CRFs, achieving comparable performance using only half the training data.
The MIT Press eBooksMarkov Random Fields for Vision and Image Processing
287 Citations2011
This volume demonstrates the power of the Markov random field in vision, treating the MRF both as a tool for modeling image data and, utilizing recently developed algorithms, as a means of making inferences about images.
Trust region Newton methods for large-scale logistic regression
287 Citations2007Chih‐Jen Lin, Ruby C. Weng +1 more
This paper applies a trust region Newton method to maximize the log-likelihood of the logistic regression model, which uses only approximate Newton steps in the beginning, but achieves fast convergence in the end.
Accelerated training of conditional random fields with stochastic gradient methods
284 Citations2006S. V. N. Vishwanathan, Nicol N. Schraudolph +2 more
Stochastic Meta-Descent (SMD), a stochastic gradient optimization method with gain vector adaptation, is applied to the training of Conditional Random Fields (CRFs) and the resulting optimizer converges to the same quality of solution over an order of magnitude faster than limited-memory BFGS.
Parsing the wall street journal using a Lexical-Functional Grammar and discriminative estimation techniques
280 Citations2001Stefan Riezler, Tracy Holloway King +4 more
The model combines full and partial parsing techniques to reach full grammar coverage on unseen data, and on a gold standard of manually annotated f-structures for a subset of the WSJ treebank, reaches 79% F-score.
Parsing the WSJ using CCG and log-linear models
279 Citations2004Stephen Clark, James Curran
A parallel implementation of the L-BFGS optimisation algorithm is described, which runs on a Beowulf cluster allowing the complete Penn Treebank to be used for estimation and a new efficient parsing algorithm for CCG which maximises expected recall of dependencies is developed.
ScholarWorks@UMassAmherst (University of Massachusetts Amherst)Accurate Information Extraction from Research Papers using Conditional Random Fields
272 Citations2004Fuchun Peng, Andrew McCallum
New state-of-the-art performance is achieved on a standard benchmark data set, reducing error in average F1 by 36%, and word error rate by 78% in comparison with the previous best SVM results.
Name Tagging with Word Clusters and Discriminative Training
270 Citations2004S.L. Miller, Jethran Guinness +1 more
A technique for augmenting annotated training data with hierarchical word clusters that are automatically derived from a large unannotated corpus that achieves a 25% reduction in error over the state-of-the-art HMM trained on the same material.
arXiv (Cornell University)MCMC for doubly-intractable distributions
260 Citations2012Iain Murray, Zoubin Ghahramani +1 more
now publishers, Inc. eBooksStructured Learning and Prediction in Computer Vision
255 Citations2010Sebastian Nowozin
Discriminative training of Markov logic networks
248 Citations2005Parag Singla, Pedro Domingos
This paper extends Collins’s (2002) voted perceptron algorithm for HMMs to MLNs by replacing the Viterbi algorithm with a weighted satisfiability solver, and proposes a discriminative approach to training MLNs.
IEEE Transactions on Information TheoryTree-based reparameterization framework for analysis of sum-product and related algorithms
238 Citations2003Martin J. Wainwright, Tommi Jaakkola +1 more
A tree-based reparameterization (TRP) framework is presented that provides a new conceptual view of a large class of algorithms for computing approximate marginals in graphs with cycles, which includes the belief propagation or sum-product algorithm as well as variations and extensions of BP.
Residual Belief Propagation: Informed Scheduling for Asynchronous Message Passing
237 Citations2012Gal Elidan, Ian McGraw +1 more
Conditional Models of Identity Uncertainty with Application to Noun Coreference
233 Citations2004Andrew McCallum, Ben Wellner
Several discriminative, conditional-probability models for coreference analysis are introduced, all examples of undirected graphical models that can incorporate a great variety of features of the input without having to be concerned about their dependencies.
Named entity recognition with a maximum entropy approach
227 Citations2003Hai Leong Chieu, Hwee Tou Ng
The named entity recognition (NER) task involves identifying noun phrases that are names, and assigning a class to each name.
ScholarWorks@UMassAmherst (University of Massachusetts Amherst)Extracting Social Networks and Contact Information From Email and the Web
225 Citations2004Aron Culotta, Ron Bekkerman +1 more
An end-to-end system that extracts a user's social network and its members' contact information given the user's email inbox and discusses the capabilities of the system for address book population, expert-finding, and social network analysis.
Painless Unsupervised Learning with Features
224 Citations2010Taylor Berg-Kirkpatrick, Alexandre Bouchard‐Côté +2 more
This work shows how features can easily be added to standard generative models for unsupervised learning, without requiring complex new training methods, and applies this technique to part-of-speech induction, grammar induction, word alignment, and word segmentation.
arXiv (Cornell University)Slow Learners are Fast
202 Citations2009John Langford, Alexander J. Smola +1 more
This paper proves that online learning with delayed updates converges well, thereby facilitating parallel online learning.
FigshareDiscriminative Fields for Modeling Spatial Dependencies in Natural Images
198 Citations2018Sanjiv Kumar, Martial Hebert
The proposed DRF model exploits local discriminative models and allows to relax the assumption of conditional independence of the observed data given the labels, commonly used in the Markov Random Field (MRF) framework.
Empirical Methods in Natural Language ProcessingMax-Margin Parsing
198 Citations2004Ben Taskar, Dan Klein +3 more
A novel discriminative approach to parsing inspired by the large-margin criterion underlying support vector machines is presented, which allows one to efficiently learn a model which discriminates among the entire space of parse trees, as opposed to reranking the top few candidates.
ScholarWorks@UMassAmherst (University of Massachusetts Amherst)FACTORIE: Probabilistic Programming via Imperatively Defined Factor Graphs
198 Citations2009Andrew McCallum, Karl Schultz +1 more
This work advocates using an imperative language to express various aspects of model structure, inference, and learning, and implements imperatively defined factor graphs in a system called FACTORIE, a software library for an object-oriented, strongly-typed, functional language.
…
