Edinburgh Research Explorer (University of Edinburgh)Open access
Charles Sutton, Andrew McCallum
Citations711
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
Abstract
4.1 Linear-Chain CRFs 4.2 Inference in Graphical Models 4.
Keywords
Computer ScienceDecision SciencesEngineering
Proceedings of the IEEEGradient-based learning applied to document recognition
58,219 Citations1998Yann LeCun, Léon Bottou +2 more
This paper reviews various methods applied to handwritten character recognition and compares them on a standard handwritten digit recognition task, and Convolutional neural networks are shown to outperform all other techniques.
Probabilistic Reasoning in Intelligent Systems: Networks of Plausible Inference
16,927 Citations1988Judea Pearl
The author provides a coherent explication of probability as a language for reasoning with partial belief and offers a unifying perspective on other AI approaches to uncertainty, such as the Dempster-Shafer formalism, truth maintenance systems, and nonmonotonic logic.
Object recognition from local scale-invariant features
16,195 Citations1999David Lowe
Experimental results show that robust object recognition can be achieved in cluttered partially occluded images with a computation time of under 2 seconds.
ScholarlyCommons (University of Pennsylvania)Conditional Random Fields: Probabilistic Models for Segmenting and Labeling Sequence Data
12,978 Citations2001John Lafferty, Andrew McCallum +1 more
This work presents iterative parameter estimation algorithms for conditional random fields and compares the performance of the resulting models to HMMs and MEMMs on synthetic and natural-language data.
IEEE Transactions on Pattern Analysis and Machine IntelligenceObject Detection with Discriminatively Trained Part-Based Models
9,976 Citations2009Pedro F. Felzenszwalb, Ross Girshick +2 more
An object detection system based on mixtures of multiscale deformable part models that is able to represent highly variable object classes and achieves state-of-the-art results in the PASCAL object detection challenges is described.
The Annals of Mathematical StatisticsA Stochastic Approximation Method
9,581 Citations1951Herbert Robbins, Sutton Monro
Numerical Optimization
9,247 Citations1999Jorge Nocedal, Stephen J. Wright
Building a Large Annotated Corpus of English: The Penn Treebank
7,528 Citations1993Mitchell P. Marcus
IEEE Transactions on Information TheoryFactor graphs and the sum-product algorithm
6,481 Citations2001Frank R. Kschischang, Brendan J. Frey +1 more
A generic message-passing algorithm, the sum-product algorithm, that operates in a factor graph, that computes-either exactly or approximately-various marginal functions derived from the global function.
Probabilistic graphical models : principles and techniques
6,434 Citations2009Daniel L. Koller, Nir Friedman
The framework of probabilistic graphical models, presented in this book, provides a general approach for causal reasoning and decision making under uncertainty, allowing interpretable models to be constructed and then manipulated by reasoning algorithms.
Neural ComputationTraining Products of Experts by Minimizing Contrastive Divergence
5,017 Citations2002Geoffrey E. Hinton
A product of experts (PoE) is an interesting candidate for a perceptual system in which rapid inference is vital and generation is unnecessary because it is hard even to approximate the derivatives of the renormalization term in the combination rule.
Incorporating non-local information into information extraction systems by Gibbs sampling
3,035 Citations2005Jenny Rose Finkel, Trond Grenager +1 more
By using simulated annealing in place of Viterbi decoding in sequence models such as HMMs, CMMs, and CRFs, it is possible to incorporate non-local structure while preserving tractable inference.
Springer texts in statisticsMonte Carlo Statistical Methods
2,245 Citations1999Christian P. Robert, George Casella
The Annals of Mathematical StatisticsStochastic Estimation of the Maximum of a Regression Function
2,143 Citations1952J. Kiefer, J. Wolfowitz
On Discriminative vs. Generative Classifiers: A comparison of logistic regression and naive Bayes
1,887 Citations2001Andrew Y. Ng, Michael I. Jordan
It is shown, contrary to a widely-held belief that discriminative classifiers are almost always to be preferred, that there can often be two distinct regimes of performance as the training set size is increased, one in which each algorithm does better.
Foundations and Trends® in Machine LearningGraphical Models, Exponential Families, and Variational Inference
1,780 Citations2008Martin J. Wainwright, Michael I. Jordan
The variational approach provides a complementary alternative to Markov chain Monte Carlo as a general source of approximation methods for inference in large-scale statistical models.
IEEE Transactions on Information TheoryConstructing Free-Energy Approximations and Generalized Belief Propagation Algorithms
1,627 Citations2005Jonathan S. Yedidia, William T. Freeman +1 more
This work explains how to obtain region-based free energy approximations that improve the Bethe approximation, and corresponding generalized belief propagation (GBP) algorithms, and describes empirical results showing that GBP can significantly outperform BP.
Journal of the Royal Statistical Society Series D (The Statistician)Statistical Analysis of Non-Lattice Data
1,575 Citations1975Julian Besag
arXiv (Cornell University)Introduction to the CoNLL-2002 Shared Task: Language-Independent Named Entity Recognition
1,574 Citations2002Erik F. Tjong Kim Sang
Computer science workbenchMarkov Random Field Modeling in Image Analysis
1,542 Citations2001Stan Z. Li
This detailed and thoroughly enhanced third edition presents a comprehensive study / reference to theories, methodologies and recent developments in solving computer vision problems based on MRFs, statistics and optimisation.
Mathematical ProgrammingPegasos: primal estimated sub-gradient solver for SVM
1,524 Citations2010Shai Shalev‐Shwartz, Yoram Singer +2 more
Journal of the American Statistical AssociationSampling-Based Approaches to Calculating Marginal Densities
1,519 Citations1990Alan E. Gelfand, A. F. M. Smith
Maximum Entropy Markov Models for Information Extraction and Segmentation
1,333 Citations2000Andrew McCallum, Dayne Freitag +1 more
A new Markovian sequence model is presented that allows observations to be represented as arbitrary overlapping features (such as word, capitalization, formatting, part-of-speech), and defines the conditional probability of state sequences given observation sequences.
Support vector machine learning for interdependent and structured output spaces
1,260 Citations2004Ioannis Tsochantaridis, Thomas Hofmann +2 more
This paper proposes to generalize multiclass Support Vector Machine learning in a formulation that involves features extracted jointly from inputs and outputs, and demonstrates the versatility and effectiveness of the method on problems ranging from supervised grammar learning and named-entity recognition, to taxonomic text classification and sequence alignment.
The MIT Press eBooksMap-Reduce for Machine Learning on Multicore
1,253 Citations2007Cheng-Tao Chu, Sang Kyun Kim +5 more
This work shows that algorithms that fit the Statistical Query model can be written in a certain "summation form," which allows them to be easily parallelized on multicore computers and shows basically linear speedup with an increasing number of processors.
Shallow parsing with conditional random fields
1,251 Citations2003Fei Sha, Fernando Pereira
This work shows how to train a conditional random field to achieve performance as good as any reported base noun-phrase chunking method on the CoNLL task, and better than any reported single model.
Max-Margin Markov Networks
1,251 Citations2003Ben Taskar, Carlos Guestrin +1 more
Maximum margin Markov (M3) networks incorporate both kernels, which efficiently deal with high-dimensional features, and the ability to capture correlations in structured data, and a new theoretical bound for generalization in structured domains is provided.
Early results for named entity recognition with conditional random fields, feature induction and web-enhanced lexicons
1,161 Citations2003Andrew McCallum, Wei Li
This work has shown that conditionally-trained models, such as conditional maximum entropy models, handle inter-dependent features of greedy sequence modeling in NLP well.
IEEE Journal on Selected Areas in CommunicationsTurbo decoding as an instance of Pearl's "belief propagation" algorithm
905 Citations1998Robert J. McEliece, David Mackay +1 more
It is shown that Pearl's algorithm can be used to routinely derive previously known iterative, but suboptimal, decoding algorithms for a number of other error-control systems, including Gallager's low-density parity-check codes, serially concatenated codes, and product codes.
Journal of Machine Learning ResearchOnline Passive-Aggressive Algorithms
899 Citations2006CrammerKoby, DekelOfer +3 more
Maximum mutual information estimation of hidden Markov model parameters for speech recognition
821 Citations2005L.R. Bahl, Peter Brown +2 more
A method for estimating the parameters of hidden Markov models of speech is described and recognition results are presented comparing this method with maximum likelihood estimation.
A Tutorial on Energy-Based Learning
793 Citations2006Yann LeCun, Sumit Chopra +3 more
The EBM approach provides a common theoretical framework for many learning models, including traditional discr iminative and generative approaches, as well as graph-transformer networks, co nditional random fields, maximum margin Markov networks, and several manifold learning methods.
IEEE Transactions on Information TheoryThe generalized distributive law
754 Citations2000S.M. Aji, Robert J. McEliece
Although this algorithm is guaranteed to give exact answers only in certain cases (the "junction tree" condition), unfortunately not including the cases of GTW with cycles or turbo decoding, there is much experimental evidence, and a few theorems, suggesting that it often works approximately even when it is not supposed to.
Mathematical ProgrammingRepresentations of quasi-Newton matrices and their use in limited memory methods
746 Citations1994Richard H. Byrd, Jorge Nocedal +1 more
This work derives compact representations of BFGS and symmetric rank-one matrices for optimization and presents a compact representation of the matrices generated by Broyden's update for solving systems of nonlinear equations.
Applying Conditional Random Fields to Japanese Morphological Analysis
723 Citations2004Taku Kudo, Kaoru Yamamoto +1 more
This paper shows how CRFs can be applied to situations where word boundary ambiguity exists, and confirms that CRFs offer a solution to the long-standing problems in corpus-based or statistical Japanese morphological analysis.
A comparison of algorithms for maximum entropy parameter estimation
667 Citations2002Robert Malouf
A number of algorithms for estimating the parameters of ME models are considered, including iterative scaling, gradient ascent, conjugate gradient, and variable metric methods.
Discriminative probabilistic models for relational data
637 Citations2002Ben Taskar, Pieter Abbeel +1 more
Maximum margin planning
630 Citations2006Nathan Ratliff, J. Andrew Bagnell +1 more
This work learns mappings from features to cost so an optimal policy in an MDP with these cost mimics the expert's behavior, and demonstrates a simple, provably efficient approach to structured maximum margin learning, based on the subgradient method, that leverages existing fast algorithms for inference.
Learning structural SVMs with latent variables
627 Citations2009Chun-Nam Yu, Thorsten Joachims
A large-margin formulation and algorithm for structured output prediction that allows the use of latent variables and the generality and performance of the approach is demonstrated through three applications including motiffinding, noun-phrase coreference resolution, and optimizing precision at k in information retrieval.
Semi-Markov Conditional Random Fields for Information Extraction
617 Citations2004Sunita Sarawagi, William W. Cohen
Intuitively, a semi-CRF on an input sequence x outputs a "segmentation" of x, in which labels are assigned to segments rather than to individual elements of xi, and transitions within a segment can be non-Markovian.
Introduction to the bio-entity recognition task at JNLPBA
579 Citations2004Jin-Dong Kim, Tomoko Ohta +3 more
The JNLPBA shared task of bio-entity recognition using an extended version of the GENIA version 3 named entity corpus of MEDLINE abstracts is described and a general discussion of the approaches taken by participating systems is presented.
BMC BioinformaticsOverview of BioCreAtIvE: critical assessment of information extraction for biology
542 Citations2005Lynette Hirschman, Alexander Yeh +2 more
The first BioCreAtIvE assessment provided state-of-the-art performance results for a basic task (gene name finding and normalization), where the best systems achieved a balanced 80% precision / recall or better, which potentially makes them suitable for real applications in biology.
Hidden Markov support vector machines
469 Citations2003Yasemin Altün, Ioannis Tsochantaridis +1 more
This paper presents a novel discriminative learning technique for label sequences based on a combination of the two most successful learning algorithms, Support Vector Machines and Hidden Markov Models which it is called HM-SVMs and handles dependencies between neighboring labels using Viterbi decoding.
Chinese segmentation and new word detection using conditional random fields
468 Citations2004Fuchun Peng, Fangfang Feng +1 more
The ability of linear-chain conditional random fields (CRFs) to perform robust and accurate Chinese word segmentation by providing a principled framework that easily supports the integration of domain knowledge in the form of multiple lexicons of characters and words is demonstrated.
Computer applications in the biosciencesABNER: an open source tool for automatically tagging genes, proteins and other entity names in text
467 Citations2005Burr Settles
Divergence measures and message passing
460 Citations2005Thomas P. Minka
This paper presents a unifying view of messagepassing algorithms, as methods to approximate a complex Bayesian network by a simpler network with minimum information divergence.
Machine LearningSearch-based structured prediction
436 Citations2009Hal Daumé, John Langford +1 more
Searn is an algorithm for integrating search and learning to solve complex structured prediction problems such as those that occur in natural language, speech, computational biology, and vision and comes with a strong, natural theoretical guarantee: good performance on the derived classification problems implies goodperformance on the structured prediction problem.
ACM eBooksGrabCut: Interactive Foreground Extraction Using Iterated Graph Cuts
434 Citations2023Carsten Rother, Vladimir Kolmogorov +1 more
A more powerful, iterative version of the optimisation of the graph-cut approach is developed and the power of the iterative algorithm is used to simplify substantially the user interaction needed for a given quality of result.
International Journal of Computer VisionDiscriminative Random Fields
392 Citations2006Sanjiv Kumar, Martial Hebert
This work presents Discriminative Random Fields (DRFs) to model spatial interactions in images in a discriminative framework based on the concept of Conditional Random Fields proposed by lafferty et al.(2001).
Identifying sources of opinions with conditional random fields and extraction patterns
366 Citations2005Yejin Choi, Claire Cardie +2 more
This work adopts a hybrid approach that combines Conditional Random Fields (Lafferty et al., 2001) and a variation of AutoSlog (Riloff, 1996a), and shows that the combination of these two methods performs better than either one alone.
arXiv (Cornell University)Efficiently Inducing Features of Conditional Random Fields
366 Citations2012Andrew McCallum
Conditional Random Fields for Object Recognition
338 Citations2004Ariadna Quattoni, Michael Collins +1 more
An extension of the CRF framework that incorporates hidden variables and combines class conditional CRFs into a unified framework for part-based object recognition is proposed, which allows the assumption of conditional independence of the observed data to be relaxed.
ArXiv.orgAn Ontology for CoNLL-RDF: Formal Data Structures for TSV Formats in Language Technology
332 Citations2000Sang, Erik F. Tjong Kim, Buchholz, Sabine
The CoNLL-2000 shared task: dividing text into syntactically related non-overlapping groups of words, so-called text chunking is described.
Hidden conditional random fields for phone classification
302 Citations2005Asela Gunawardana, Milind Mahajan +2 more
This paper presents the results on the TIMIT phone classification task and shows that HCRFs outperforms comparable ML and CML/MMI trained HMMs and has the ability to handle complex features without any change in training procedure.
Practical Very Large Scale CRFs
299 Citations2010Thomas Lavergne, Olivier Cappé +1 more
This paper addresses the issue of training very large CRFs, containing up to hundreds output labels and several billion features, and indicates that efficiency stems here from the sparsity induced by the use of a l penalty term.
Edinburgh Research Explorer (University of Edinburgh)Dynamic Conditional Random Fields: Factorized Probabilistic Models for Labeling and Segmenting Sequence Data
292 Citations2007Charles Sutton, Andrew McCallum +1 more
On a natural-language chunking task, it is shown that a DCRF performs better than a series of linear-chain CRFs, achieving comparable performance using only half the training data.
The MIT Press eBooksMarkov Random Fields for Vision and Image Processing
287 Citations2011
This volume demonstrates the power of the Markov random field in vision, treating the MRF both as a tool for modeling image data and, utilizing recently developed algorithms, as a means of making inferences about images.
Trust region Newton methods for large-scale logistic regression
287 Citations2007Chih‐Jen Lin, Ruby C. Weng +1 more
This paper applies a trust region Newton method to maximize the log-likelihood of the logistic regression model, which uses only approximate Newton steps in the beginning, but achieves fast convergence in the end.
Parsing the wall street journal using a Lexical-Functional Grammar and discriminative estimation techniques
280 Citations2001Stefan Riezler, Tracy Holloway King +4 more
The model combines full and partial parsing techniques to reach full grammar coverage on unseen data, and on a gold standard of manually annotated f-structures for a subset of the WSJ treebank, reaches 79% F-score.
Parsing the WSJ using CCG and log-linear models
279 Citations2004Stephen Clark, James Curran
A parallel implementation of the L-BFGS optimisation algorithm is described, which runs on a Beowulf cluster allowing the complete Penn Treebank to be used for estimation and a new efficient parsing algorithm for CCG which maximises expected recall of dependencies is developed.
ScholarWorks@UMassAmherst (University of Massachusetts Amherst)Accurate Information Extraction from Research Papers using Conditional Random Fields
272 Citations2004Fuchun Peng, Andrew McCallum
New state-of-the-art performance is achieved on a standard benchmark data set, reducing error in average F1 by 36%, and word error rate by 78% in comparison with the previous best SVM results.
Name Tagging with Word Clusters and Discriminative Training
270 Citations2004S.L. Miller, Jethran Guinness +1 more
A technique for augmenting annotated training data with hierarchical word clusters that are automatically derived from a large unannotated corpus that achieves a 25% reduction in error over the state-of-the-art HMM trained on the same material.
Discriminative training of Markov logic networks
248 Citations2005Parag Singla, Pedro Domingos
This paper extends Collins’s (2002) voted perceptron algorithm for HMMs to MLNs by replacing the Viterbi algorithm with a weighted satisfiability solver, and proposes a discriminative approach to training MLNs.
IEEE Transactions on Information TheoryTree-based reparameterization framework for analysis of sum-product and related algorithms
238 Citations2003Martin J. Wainwright, Tommi Jaakkola +1 more
A tree-based reparameterization (TRP) framework is presented that provides a new conceptual view of a large class of algorithms for computing approximate marginals in graphs with cycles, which includes the belief propagation or sum-product algorithm as well as variations and extensions of BP.
Residual Belief Propagation: Informed Scheduling for Asynchronous Message Passing
237 Citations2012Gal Elidan, Ian McGraw +1 more
Conditional Models of Identity Uncertainty with Application to Noun Coreference
233 Citations2004Andrew McCallum, Ben Wellner
Several discriminative, conditional-probability models for coreference analysis are introduced, all examples of undirected graphical models that can incorporate a great variety of features of the input without having to be concerned about their dependencies.
Named entity recognition with a maximum entropy approach
227 Citations2003Hai Leong Chieu, Hwee Tou Ng
The named entity recognition (NER) task involves identifying noun phrases that are names, and assigning a class to each name.
ScholarWorks@UMassAmherst (University of Massachusetts Amherst)Extracting Social Networks and Contact Information From Email and the Web
225 Citations2004Aron Culotta, Ron Bekkerman +1 more
An end-to-end system that extracts a user's social network and its members' contact information given the user's email inbox and discusses the capabilities of the system for address book population, expert-finding, and social network analysis.
Painless Unsupervised Learning with Features
224 Citations2010Taylor Berg-Kirkpatrick, Alexandre Bouchard‐Côté +2 more
This work shows how features can easily be added to standard generative models for unsupervised learning, without requiring complex new training methods, and applies this technique to part-of-speech induction, grammar induction, word alignment, and word segmentation.
Lecture notes in computer scienceUltraconservative Online Algorithms for Multiclass Problems
221 Citations2001Koby Crammer, Yoram Singer
This paper designs and analyzes online classification algorithms for multiclass problems in the mistake bound model and describes a family of additive ultraconservative algorithms where each algorithm in the family updates its prototypes by finding a feasible solution for a set of linear constraints that depend on the instantaneous similarity-scores.
Robust higher order potentials for enforcing label consistency
202 Citations2008Pushmeet Kohli, Ľubor Ladický +1 more
Journal of Machine Learning ResearchPosterior Regularization for Structured Latent Variable Models
202 Citations2010GanchevKuzman, GraçaJoão +2 more
This work presents an efficient algorithm for learning with posterior regularization and illustrates its versatility on a diverse set of structural constraints such as bijectivity, symmetry and group sparsity in several large scale experiments, including multi-view learning, cross-lingual dependency grammar induction, unsupervised part-of-speech induction, and bitext word alignment.
FigshareDiscriminative Fields for Modeling Spatial Dependencies in Natural Images
198 Citations2018Sanjiv Kumar, Martial Hebert
The proposed DRF model exploits local discriminative models and allows to relax the assumption of conditional independence of the observed data given the labels, commonly used in the Markov Random Field (MRF) framework.
ScholarWorks@UMassAmherst (University of Massachusetts Amherst)FACTORIE: Probabilistic Programming via Imperatively Defined Factor Graphs
198 Citations2009Andrew McCallum, Karl Schultz +1 more
This work advocates using an imperative language to express various aspects of model structure, inference, and learning, and implements imperatively defined factor graphs in a system called FACTORIE, a software library for an object-oriented, strongly-typed, functional language.
Lecture notes in computer scienceLocalizing Objects While Learning Their Appearance
196 Citations2010Thomas Deselaers, Bogdan Alexe +1 more
This work proposes a conditional random field that starts from generic knowledge and then progressively adapts to the new class to enable any state-of-the-art object detector in a weakly supervised fashion, although it would normally require object location annotations.
Integer linear programming inference for conditional random fields
187 Citations2005Dan Roth, Wen-tau Yih
A novel inference procedure based on integer linear programming (ILP) and extends CRF models to naturally and efficiently support general constraint structures is proposed and Experimental evidence is supplied in the context of an important NLP problem, semantic role labeling.
Comparisons of sequence labeling algorithms and extensions
182 Citations2007Nam Hoang Nguyen, Yunsong Guo
This paper surveys the current state-of-art models for structured learning problems, including Hidden Markov Model (HMM), Conditional Random Fields (CRF), Averaged Perceptron (AP), Structured SVMs (SVMstruct), Max Margin Markov Networks (M3N), and an integration of search and learning algorithm (SEARN).
Efficient, Feature-based, Conditional Random Field Parsing
180 Citations2008Jenny Rose Finkel, Alex Kleeman +1 more
This work presents the first general, featurerich discriminative parser, based on a conditional random field model, which has been successfully scaled to the full WSJ parsing data, and achieves state-of-the-art results.
A discriminative matching approach to word alignment
166 Citations2005Ben Taskar, Simon Lacoste-Julien +1 more
This work presents a discriminative, large-margin approach to feature-based matching for word alignment, which achieves AER performance close to IBM Model 4, in much less time.
CORE Scholar (Wright State University)Collective Segmentation and Labeling of Distant Entities in Information Extraction
148 Citations2004Charles Sutton, Andrew McCallum
This work presents a CRF that explicitly represents dependencies between the labels of pairs of similar words in a document, and shows that learning these dependencies leads to a 13.7% reduction in error on the field that had caused the most repetition errors.
arXiv (Cornell University)Piecewise Training for Undirected Models
146 Citations2012Charles Sutton, Andrew McCallum
Collective information extraction with relational Markov networks
139 Citations2004Răzvan Bunescu, Raymond J. Mooney
A new IE method is presented that employs Relational Markov Networks (a generalization of CRFs), which can represent arbitrary dependencies between extractions, which allows for "collective information extraction" that exploits the mutual influence between possible extractions.
Semi-supervised conditional random fields for improved sequence segmentation and labeling
130 Citations2006Feng Jiao, Shaojun Wang +3 more
A new semi-supervised training procedure for conditional random fields (CRFs) that can be used to train sequence segmentors and labelers from a combination of labeled and unlabeled training data is presented, based on extending the minimum entropy regularization framework to the structured prediction case.
CORE Scholar (Wright State University)Generalized Expectation Criteria for Semi-Supervised Learning of Conditional Random Fields
130 Citations2008Gideon Mann, Andrew McCallum
This paper presents a semi-supervised training method for linear-chain conditional random fields that makes use of labeled features rather than labeled instances by using generalized expectation criteria to express a preference for parameter settings in which the model’s distribution on unlabeled data matches a target distribution.
Efficient Training of Conditional Random Fields
128 Citations2002Hanna Wallach
This thesis explores a number of parameter estimation techniques for conditional random fields, a recently introduced probabilistic model for labelling and segmenting sequential data, and hypothesises that general numerical optimisation techniques result in improved performance over iterative scaling algorithms for training CRFs.
Structured Learning with Approximate Inference
128 Citations2007Alex Kulesza, Fernando Pereira
It is shown in particular that learning can fail even with an approximate inference method with rigorous approximation guarantees, and argued that without understanding combinations of inference and learning, such as these that are appropriately compatible, learning performance under approximate inference cannot be guaranteed.
An asymptotic analysis of generative, discriminative, and pseudolikelihood estimators
127 Citations2008Percy Liang, Michael I. Jordan
This paper presents a unified framework for studying parameter estimators, which allows them to compare their relative (statistical) efficiencies, and suggests that modeling more of the data tends to reduce variance, but at the cost of being more sensitive to model misspecification.
Confidence estimation for information extraction
115 Citations2004Aron Culotta, Andrew McCallum
This work evaluates a information extraction system based on a linear-chain conditional random field (CRF), a probabilistic model which has performed well on information extraction tasks because of its ability to capture arbitrary, overlapping features of the input in a Markov model.
A Conditional Random Field for Discriminatively-Trained Finite-State String Edit Distance
114 Citations2005Andrew McCallum, Kedar Bellare +1 more
This paper presents discriminative string-edit CRFs, a finite-state conditional random field model for edit sequences between strings, trained on both positive and negative instances of string pairs.
Learning from measurements in exponential families
114 Citations2009Percy Liang, Michael I. Jordan +1 more
A Bayesian decision-theoretic framework is presented, which allows us to both integrate diverse measurements and choose new measurements to make, and a variational inference algorithm is used, which exploits exponential family duality.
Discriminative word alignment with conditional random fields
106 Citations2006Phil Blunsom, Trevor Cohn
A novel approach for inducing word alignments from sentence aligned data using a Conditional Random Field, a discriminative model, which is estimated on a small supervised training set, and which has efficient training and decoding processes which both find globally optimal solutions.
On the Convergence Properties of Contrastive Divergence
91 Citations2010Ilya Sutskever, Tijmen Tieleman
This paper analyzes the CD1 update rule for Restricted Boltzmann Machines with binary variables, and shows that the regularized CD update has a fixed point for a large class of regularization functions using Brower’s fixed point theorem.
PLoS Computational BiologyGlobal Discriminative Learning for Higher-Accuracy Computational Gene Prediction
90 Citations2007Axel Bernal, Koby Crammer +2 more
CRAIG, a new program for ab initio gene prediction based on a conditional random field model with semi-Markov structure that is trained with an online large-margin algorithm related to multiclass SVMs, shows significant improvements in prediction accuracy over published gene predictors that use intrinsic features only.
BioinformaticsRNA secondary structural alignment with conditional random fields
85 Citations2005Kengo Sato, Yasubumi Sakakibara
Experimental results clearly show that the parameter estimation with CRFs can outperform all the other existing methods for structural alignments of RNA sequences and structural alignment search based on CRFs is more accurate for predicting non-coding RNA regions than the other scoring methods.
Edinburgh Research Archive (University of Edinburgh)MCMC for doubly-intractable distributions
84 Citations2012Iain Murray, Zoubin Ghahramani +1 more
This paper provides a generalization of M0ller et al. (2004) and a new MCMC algorithm, which obtains better acceptance probabilities for the same amount of exact sampling, and removes the need to estimate model parameters before sampling begins.
Learning to extract information from semi-structured text using a discriminative context free grammar
75 Citations2005Paul Viola, Mukund Narasimhan
It is shown that a statistical parsing approach results in a 50% reduction in error rate and this system also has the advantage of being interactive, similar to the system described in [9].
An Empirical Comparison of Supervised Learning Algorithms Using Different Performance Metrics
75 Citations2005Rich Caruana, Alex Niculescu-Mizil
A large-scale empirical comparison between ten learning methods finds that the best models are boosted trees, random forests, and unscaled neural nets.
Advances in Markov chain Monte Carlo methods
73 Citations2007Iain Murray
This thesis proposes and investigates several new Monte Carlo algorithms, both for evaluating normalizing constants and for improved sampling of distributions, and develops novel exact-sampling-based MCMC methods, the Exchange Algorithm and Latent Histories.
…
