Introduction to Semi-Supervised Learning
Synthesis lectures on artificial intelligence and machine learningPublished 1 January 2009
Xiaojin Zhu, Andrew B. Goldberg
Citations1,804
SJR quartileQ4
SJR score0.23
SNIP2.20
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
This introductory book presents some popular semi-supervised learning models, including self-training, mixture models, co-training and multiview learning, graph-based methods, and semi- supervised support vector machines, and discusses their basic mathematical formulation.
Abstract
Semi-supervised learning is a learning paradigm concerned with the study of how computers and natural systems such as humans learn in the presence of both labeled and unlabeled data. Traditionally, le
Keywords
Computer Science
Journal of the Royal Statistical Society Series B (Statistical Methodology)Maximum Likelihood from Incomplete Data Via the <i>EM</i> Algorithm
49,657 Citations1977A. P. Dempster, N. M. Laird +1 more
TechnometricsStatistical Learning Theory
26,913 Citations1999Yuhai Wu, Vladimir Vapnik
Presenting a method for determining the necessary and sufficient conditions for consistency of learning process, the author covers function estimates from small data pools, applying these estimations to real-life problems, and much more.
Springer series in statisticsThe Elements of Statistical Learning
24,344 Citations2001Trevor Hastie, J. Friedman +1 more
Choice Reviews OnlineArtificial intelligence: a modern approach
22,205 Citations1995Stuart Russell, Peter Norvig +2 more
Journal of Electronic ImagingPattern Recognition and Machine Learning
21,976 Citations2007Christopher Bishop
Probability Distributions, linear models for Regression, Linear Models for Classification, Neural Networks, Graphical Models, Mixture Models and EM, Sampling Methods, Continuous Latent Variables, Sequential Data are studied.
Cambridge University Press eBooksAn Introduction to Support Vector Machines and Other Kernel-based Learning Methods
13,883 Citations2000Nello Cristianini, John Shawe‐Taylor
Mathematical ProgrammingOn the limited memory BFGS method for large scale optimization
8,529 Citations1989Dong C. Liu, Jorge Nocedal
The numerical tests indicate that the L-BFGS method is faster than the method of Buckley and LeNir, and is better able to use additional storage to accelerate convergence, and the convergence properties are studied to prove global convergence on uniformly convex problems.
Cambridge University Press eBooksKernel Methods for Pattern Analysis
6,598 Citations2004John Shawe‐Taylor, Nello Cristianini
This book provides an easy introduction for students and researchers to the growing field of kernel-based pattern analysis, demonstrating with examples how to handcraft an algorithm or a kernel for a new specific application, and covering all the necessary conceptual and mathematical tools to do so.
Combining labeled and unlabeled data with co-training
5,604 Citations1998Avrim Blum, Tom M. Mitchell
Technical reportsMaking Large-Scale SVM Learning Practical
4,317 Citations2006Thorsten Joachims
This chapter presents algorithmic and computational results developed for SVM light V 2.0, which make large-scale SVM training more practical and give guidelines for the application of SVMs to large domains.
The MIT Press eBooksSemi-Supervised Learning
4,308 Citations2006Olivier Chapelle, Schölkopf, B. +1 more
This first comprehensive overview of semi-supervised learning presents state-of-the-art algorithms, a taxonomy of the field, selected applications, benchmark experiments, and perspectives on ongoing and future research.
A theory of the learnable
4,242 Citations1984Leslie G. Valiant
This paper regards learning as the phenomenon of knowledge acquisition in the absence of explicit programming, and gives a precise methodology for studying this phenomenon from a computational viewpoint.
Minds at UW (University of Wisconsin)Semi-Supervised Learning Literature Survey
3,868 Citations2005Xiaojin Zhu
The study clearly indicates that the common practice of stripwise precommercial thinning is unjustified, and the justification of heavy 'chessboard' thinning (with pruning) depends on whether the potential reduction in rotation length and the improvement in wood quality outweigh the discounted costs of pre-commercial thinning and selection and pruning of crop trees.
MPG.PuRe (Max Planck Society)Learning with Local and Global Consistency
3,750 Citations2003Dengyong Zhou, Olivier Bousquet +3 more
A principled approach to semi-supervised learning is to design a classifying function which is sufficiently smooth with respect to the intrinsic structure collectively revealed by known labeled and unlabeled points.
Semi-supervised learning using Gaussian fields and harmonic functions
3,377 Citations2003Xiaojin Zhu, Zoubin Ghahramani +1 more
A sentimental education
3,343 Citations2004Bo Pang, Lillian Lee
A novel machine-learning method is proposed that applies text-categorization techniques to just the subjective portions of the document, which greatly facilitates incorporation of cross-sentence contextual constraints.
Manifold Regularization: A Geometric Framework for Learning from Labeled and Unlabeled Examples
3,266 Citations2006Mikhail Belkin, Partha Niyogi +1 more
A semi-supervised framework that incorporates labeled and unlabeled data in a general-purpose learner is proposed and properties of reproducing kernel Hilbert spaces are used to prove new Representer theorems that provide theoretical basis for the algorithms.
A comparison of event models for naive bayes text classification
3,224 Citations1998Andrew McCallum, Kamal Nigam
It is found that the multi-variate Bernoulli performs well with small vocabulary sizes, but that the multinomial performs usually performs even better at larger vocabulary sizes--providing on average a 27% reduction in error over the multi -variateBernoulli model at any vocabulary size.
Machine LearningText Classification from Labeled and Unlabeled Documents using EM
2,749 Citations2000Kamal Nigam, Andrew Kachites McCallum +2 more
This paper shows that the accuracy of learned text classifiers can be improved by augmenting a small number of labeled training documents with a large pool of unlabeled documents, and presents two extensions to the algorithm that improve classification accuracy under these conditions.
Transductive Inference for Text Classification using Support Vector Machines
2,717 Citations1999Thorsten Joachims
An analysis of why TSVMs are well suited for text classi(cid:12)cation is presented, and an algorithm for training TSVMs e(cid:14)-ciently, handling 10,000 examples and more is proposed.
Unsupervised word sense disambiguation rivaling supervised methods
2,408 Citations1995David Yarowsky
An unsupervised learning algorithm for sense disambiguation that, when trained on unannotated English text, rivals the performance of supervised techniques that require time-consuming hand annotations.
Journal of Experimental Psychology GeneralAttention, similarity, and the identification-categorization relationship.
2,225 Citations1986Robert M. Nosofsky
A unified quantitative approach to modeling subjects' identification and categorization of multidimensional perceptual stimuli is proposed and tested and some support was gained for the hypothesis that subjects distribute attention among component dimensions so as to optimize categorization performance.
Labeling images with a computer game
2,222 Citations2004Luis von Ahn, Laura Dabbish
A new interactive system: a game that is fun and can be used to create valuable output that addresses the image-labeling problem and encourages people to do the work by taking advantage of their desire to be entertained.
SWITCHBOARD: telephone speech corpus for research and development
2,140 Citations1992J. Godfrey, E. Holliman +1 more
Lecture notes in computer scienceRademacher and Gaussian Complexities: Risk Bounds and Structural Results
2,111 Citations2001Peter L. Bartlett, Shahar Mendelson
This work investigates the use of certain data-dependent estimates of the complexity of a function class called Rademacher and Gaussian complexities and proves general risk bounds in terms of these complexities in a decision theoretic setting.
MPG.PuRe (Max Planck Society)Large Margin Methods for Structured and Interdependent Output Variables
1,952 Citations2005Ioannis Tsochantaridis, Thorsten Joachims +2 more
This paper proposes to appropriately generalize the well-known notion of a separation margin and derive a corresponding maximum-margin formulation and presents a cutting plane algorithm that solves the optimization problem in polynomial time for a large class of problems.
The MIT Press eBooksAn Introduction to Computational Learning Theory
1,733 Citations1994Michael Kearns, Umesh Vazirani
The probably approximately correct learning model Occam's razor the Vapnik-Chervonenkis dimension weak and strong learning learning in the presence of noise inherent unpredictability reducibility in PAC learning learning finite automata is described.
Self-taught learning
1,533 Citations2007Rajat Raina, Alexis Battle +3 more
An approach to self-taught learning that uses sparse coding to construct higher-level features using the unlabeled data to form a succinct input representation and significantly improve classification performance.
ACM Transactions on GraphicsColorization using optimization
1,439 Citations2004Anat Levin, Dani Lischinski +1 more
Journal of Machine Learning ResearchA Framework for Learning Predictive Structures from Multiple Tasks and Unlabeled Data
1,372 Citations2005Rie Kubota Ando, Tong Zhang
This paper presents a general framework in which the structural learning problem can be formulated and analyzed theoretically, and relate it to learning with unlabeled data, and algorithms for structural learning will be proposed, and computational issues will be investigated.
Max-Margin Markov Networks
1,251 Citations2003Ben Taskar, Carlos Guestrin +1 more
Maximum margin Markov (M3) networks incorporate both kernels, which efficiently deal with high-dimensional features, and the ability to capture correlations in structured data, and a new theoretical bound for generalization in structured domains is provided.
The MIT Press eBooksLearning with Hypergraphs: Clustering, Classification, and Embedding
1,240 Citations2007Dengyong Zhou, Jiayuan Huang +1 more
This paper generalizes the powerful methodology of spectral clustering which originally operates on undirected graphs to hypergraphs, and further develop algorithms for hypergraph embedding and transductive classification on the basis of the spectral hypergraph clustering approach.
IEEE Transactions on Knowledge and Data EngineeringTri-training: exploiting unlabeled data using three classifiers
1,185 Citations2005Zhi‐Hua Zhou, Ming Li
Experiments on UCI data sets and application to the Web page classification task indicate that tri-training can effectively exploit unlabeled data to enhance the learning performance.
Analyzing the effectiveness and applicability of co-training
1,028 Citations2000Kamal Nigam, Rayid Ghani
It is demonstrated that when learning from labeled and unlabeled data, algorithms explicitly leveraging a natural independent split of the features outperform algorithms that do not and may out-perform algorithms not using a split.
Semi-supervised Learning by Entropy Minimization.
1,021 Citations2005Yves Grandvalet, Yoshua Bengio
This framework, which motivates minimum entropy regularization, enables to incorporate unlabeled data in the standard supervised learning, and includes other approaches to the semi-supervised problem as particular or limiting cases.
Research Showcase @ Carnegie Mellon University (Carnegie Mellon University)Learning from Labeled and Unlabeled Data using Graph Mincuts
947 Citations2018Avrim Blum, Shuchi Chawla
An algorithm based on finding minimum cuts in graphs, that uses pairwise relationships among the examples in order to learn from both labeled and unlabeled data is considered.
Lecture notes in computer scienceKernels and Regularization on Graphs
839 Citations2003Alexander J. Smola, Risi Kondor
It is shown that the class of positive, monotonically decreasing functions on the unit interval leads to kernels and corresponding regularization operators and can be found as a special case of the reasoning.
Unsupervised Models for Named Entity Classification
817 Citations1999Michael Collins, Yoram Singer
It is shown that the use of unlabeled data can reduce the requirements for supervision to just 7 simple "seed" rules, gaining leverage from natural redundancy in the data.
Semi-Supervised Support Vector Machines
793 Citations1998Kristin P. Bennett, Ayhan Demiriz
A general S3VM model is proposed that minimizes both the misclassification error and the function capacity based on all the available data that can be converted to a mixed-integer program and then solved exactly using integer programming.
Semi-Supervised Self-Training of Object Detection Models
791 Citations2005Charles Rosenberg, Martial Hebert +1 more
The key contributions of this empirical study are to demonstrate that a model trained in this manner can achieve results comparable to a modeltrained in the traditional manner using a much larger set of fully labeled data, and that a training data selection metric that is defined independently of the detector greatly outperforms a selection metric based on the detection confidence generated by the detector.
Diffusion Kernels on Graphs and Other Discrete Input Spaces
732 Citations2002Risi Kondor, John Lafferty
This paper proposes a general method of constructing natural families of kernels over discrete structures, based on the matrix exponentiation idea, and focuses on generating kernels on graphs, for which a special class of exponential kernels called diffusion kernels are proposed.
Semi-Supervised Classification by Low Density Separation
710 Citations2005Olivier Chapelle, Alexander Zien +1 more
Three semi-supervised algorithms are proposed: deriving graph-based distances that emphazise low density regions between clusters, followed by training a standard SVM, and optimizing the Transductive SVM objective function by gradient descent.
Transductive learning via spectral graph partitioning
649 Citations2003Thorsten Joachims
This work proposes an algorithm that robustly achieves good generalization performance and that can be trained efficiently, and shows a connection to transductive Support Vector Machines, and that an effective Co-Training algorithm arises as a special case.
IEEE Transactions on Geoscience and Remote SensingThe effect of unlabeled samples in reducing the small sample size problem and mitigating the Hughes phenomenon
572 Citations1994B.M. Shahshahani, D. A. Landgrebe
By using additional unlabeled samples that are available at no extra cost, the performance may be improved, and therefore the Hughes phenomenon can be mitigated and therefore more representative estimates can be obtained.
Partially labeled classification with Markov random walks
563 Citations2001Martin Szummer, Tommi Jaakkola
This work combines a limited number of labeled examples with a Markov random walk representation over the unlabeled examples and develops and compares several estimation criteria/algorithms suited to this representation.
The Annals of StatisticsConsistency of spectral clustering
535 Citations2008Ulrike von Luxburg, Mikhail A. Belkin +1 more
It is proved that one of the two major classes of spectral clustering (normalized clustering) converges under very general conditions, while the other is only consistent under strong additional assumptions, which are not always satisfied in real data.
Partially Supervised Classification of Text Documents
516 Citations2002Bing Liu, Wee Sun Lee +2 more
This paper shows that the problem of identifying documents from a set of documents of a particular topic or class P and a large set M of mixed documents, and that under appropriate conditions, solutions to the constrained optimization problem will give good solution to the partially supervised classification problem.
Learning subjective nouns using extraction pattern bootstrapping
499 Citations2003Ellen Riloff, Janyce Wiebe +1 more
The goal of the research is to develop a system that can distinguish subjective sentences from objective sentences, and a Naive Bayes classifier is trained using the subjective nouns, discourse features, and subjectivity clues identified in prior research.
ScienceThe Learning of Categories: Parallel Brain Systems for Item Memory and Category Knowledge
487 Citations1993Barbara J. Knowlton, Larry R. Squire
Amnesic patients and control subjects performed similarly at classifying novel patterns according to whether they belonged to the same category as a set of training patterns, and the amnesics were impaired at recognizing which dot patterns had been presented for training.
Combining active learning and semi-supervised learning using Gaussian fields and harmonic functions
471 Citations2003Xiaojin Zhu, John Lafferty +1 more
The active learning scheme requires a much smaller number of queries to achieve high accuracy compared with random query selection, and the semi-supervised learning scheme requires a much smaller number of queries to achieve high accuracy compared with random query selection.
Encyclopedia of Machine Learning and Data MiningLearning from Labeled and Unlabeled Data
469 Citations2017Claude Sammut, Geoffrey I. Webb
Cluster Kernels for Semi-Supervised Learning
461 Citations2002Olivier Chapelle, Jason Weston +1 more
A framework to incorporate unlabeled data in kernel classifier, based on the idea that two points in the same cluster are more likely to have the same label is proposed by modifying the eigenspectrum of the kernel matrix.
Enhancing Supervised Learning with Unlabeled Data
459 Citations2000Sally A. Goldman, Yan Zhou
A new semi-supervised learning method called co-learning that is designed to use unlabeled data to enhance standard supervised learning algorithms to leverage off the fact that they have different representations of the hypotheses and are likely to detect different patterns in labeled data.
Beyond the point cloud
443 Citations2005Vikas Sindhwani, Partha Niyogi +1 more
This paper constructs a family of data-dependent norms on Reproducing Kernel Hilbert Spaces (RKHS) that allow the structure of the RKHS to reflect the underlying geometry of the data.
Constrained Clustering
433 Citations2008Sugato Basu, Ian Davidson +1 more
This volume delivers thorough coverage of the capabilities and limitations of constrained clustering methods as well as introduces new types of constraints and clustering algorithms.
Journal of Machine Learning ResearchOptimization Techniques for Semi-Supervised Support Vector Machines
410 Citations2008Olivier Chapelle, Vikas Sindhwani +1 more
The performance and behavior of various S3VMs algorithms is studied together, under a common experimental setting, to review key ideas in this literature on semi-supervised support Vector Machines.
Learning from labeled and unlabeled data on a directed graph
403 Citations2005Dengyong Zhou, Jiayuan Huang +1 more
A general framework for learning from labeled and unlabeled data on a directed graph in which the structure of the graph including the directionality of the edges is considered, which generalizes the spectral clustering approach for undirected graphs.
Deep learning via semi-supervised embedding
387 Citations2008Jason Weston, Frédéric Ratle +1 more
It is shown how nonlinear embedding algorithms popular for use with shallow semi-supervised learning techniques such as kernel methods can be applied to deep multilayer architectures, either as a regularizer at the output layer, or on each layer of the architecture.
Scientific Repository (Petra Christian University)Gaussian Processes for Ordinal Regression
373 Citations2005Wei Chu, Zoubin Ghahramani
Trading convexity for scalability
361 Citations2006Ronan Collobert, Fabian H. Sinz +2 more
It is shown how concave-convex programming can be applied to produce faster SVMs where training errors are no longer support vectors, and much faster Transductive SVMs.
Learning with positive and unlabeled examples using weighted logistic regression
334 Citations2003Wee Sun Lee, Bing Liu
A performance measure that can be estimated from positive and unlabeled examples for evaluating retrieval performance, which is proportional to the product of precision and recall, can be used with a validation set to select regularization parameters for logistic regression.
Brunel University Research Archive (BURA) (Brunel University London)Two view learning: SVM-2K, Theory and Practice
331 Citations2005Jason Farquhar, David R. Hardoon +3 more
Seeing stars when there aren't many stars
319 Citations2006Andrew B. Goldberg, Xiaojin Zhu
A graph-based semi-supervised learning algorithm is presented to address the sentiment analysis task of rating inference and achieves significantly better predictive accuracy over other methods that ignore the unlabeled examples during training.
Neural Information Processing SystemsA Mixture of Experts Classifier with Learning Based on Both Labelled and Unlabelled Data
289 Citations1996David J. Miller, Hasan S. Uyar
A classifier structure and learning algorithm that make effective use of unlabelled data to improve performance and is a "mixture of experts" structure that is equivalent to the radial basis function (RBF) classifier, but unlike RBFs, is amenable to likelihood-based training.
Active + Semi-supervised Learning = Robust Multi-View Learning
288 Citations2002Ion Muslea, Steven Minton +1 more
A new multi-view algorithm, Co-EMT, which combines semi-supervised and active learning is introduced, which outperforms the other algorithms both on the parameterized problems and on two additional real world domains.
The MIT Press eBooksPAC Generalization Bounds for Co-training
287 Citations2002Sanjoy Dasgupta, Michael L. Littman +1 more
A new PAC-style bound on generalization error is given which justifies both the use of confidences — partial rules and partial labeling of the unlabeled data — and theUse of an agreement-based objective function as suggested by Collins and Singer.
The MIT Press eBooksSpectral Methods for Dimensionality Reduction
271 Citations2006Lawrence K. Saul, Kilian Q. Weinberger +3 more
Semi-supervised learning using randomized mincuts
268 Citations2004Avrim Blum, John Lafferty +2 more
The experiments on several datasets show that when the structure of the graph supports small cuts, this can result in highly accurate classifiers with good accuracy/coverage tradeoffs, and can be given theoretical justification from both a Markov random field perspective and from sample complexity considerations.
Label propagation through linear neighborhoods
261 Citations2006Fei Wang, Changshui Zhang
A novel graph-based semi supervised learning approach is proposed based on a linear neighborhood model, which assumes that each data point can be linearly reconstructed from its neighborhood, and can propagate the labels from the labeled points to the whole data set using these linear neighborhoods with sufficient smoothness.
Research Showcase @ Carnegie Mellon University (Carnegie Mellon University)Co-Training and Expansion: Towards Bridging Theory and Practice
261 Citations2018Maria-Florina Balcan, Avrim Blum +1 more
A much weaker "expansion" assumption on the underlying data distribution is proposed, that is proved to be sufficient for iterative co-training to succeed given appropriately strong PAC-learning algorithms on each feature set, and that to some extent is necessary as well.
Machine LearningMaximum Entropy Discrimination
259 Citations2004Tony Jebara
A general framework for discriminative estimation based on the maximum entropy principle and its extensions is presented and preliminary experimental results are indicative of the potential in these techniques.
International Joint Conference on Artificial IntelligenceSemi-supervised regression with co-training
244 Citations2005Zhi‐Hua Zhou, Ming Li
Experiments show that COREG can effectively exploit unlabeled data to improve regression estimates and is proposed as a co-training style semi-supervised regression algorithm.
An RKHS for multi-view learning and manifold co-regularization
224 Citations2008Vikas Sindhwani, David S. Rosenberg
This paper constructs a single RKHS with a data-dependent "co-regularization" norm that reduces these approaches to standard supervised learning and proposes a co- regularization based algorithmic alternative to manifold regularization that leads to major empirical improvements on semi-supervised tasks.
Semi-supervised learning of compact document representations with deep networks
223 Citations2008Marc ' Aurelio Ranzato, Martin Szummer
An algorithm to learn text document representations based on semi-supervised autoencoders that are stacked to form a deep network that can be trained efficiently on partially labeled corpora, producing very compact representations of documents, while retaining as much class information and joint word statistics as possible.
Journal of Machine Learning ResearchGraph Laplacians and their Convergence on Random Neighborhood Graphs
217 Citations2007Matthias Hein, Jean-Yves Audibert +1 more
This paper determines the pointwise limit of three different graph Laplacians used in the literature as the sample size increases and the neighborhood size approaches zero and shows that for a uniform measure on the submanifold all graph LaPLacians have the same limit up to constants.
Inference with the Universum
208 Citations2006Jason Weston, Ronan Collobert +3 more
An algorithm is described to leverage the Universum by maximizing the number of observed contradictions, and it is shown experimentally that this approach delivers accuracy improvements over using labeled data alone.
Optimization methods & softwareSemi-superyised support vector machines for unlabeled data classification
204 Citations2001Glenn Fung, O. L. Mangasarian
Learning Classification with Unlabeled Data
204 Citations1993Virginia R. de
This paper shows that minimizing the disagreement between the outputs of networks processing patterns from these different modalities is a sensible approximation to minimizing the number of misclassifications in each modality, and leads to similar results.
Perception & PsychophysicsOn the dominance of unidimensional rules in unsupervised categorization
203 Citations1999F. Gregory Ashby, Sarah Queller +1 more
In several experiments, observers tried to categorize stimuli constructed from two separable stimulus dimensions in the absence of any trial-by-trial feedback to support the hypothesis that people are constrained to use unidimensional rules.
Pattern Recognition LettersOn the exponential value of labeled samples
201 Citations1995Vittorio Castelli, Thomas M. Cover
The first labeled sample reduces the risk from 1 2 to 2R ∗ (1−R∗ ) and subsequent labeled samples in the training set reduce the probability of error exponentially fast to the Bayes risk.
Unlabeled data: Now it helps, now it doesn't
195 Citations2008Aarti Singh, Robert D. Nowak +1 more
A finite sample analysis is developed that characterizes the value of un-labeled data and quantifies the performance improvement of SSL compared to supervised learning, and shows that there are large classes of problems for which SSL can significantly outperform supervised learning in finite sample regimes and sometimes also in terms of error convergence rates.
Semi-supervised learning of mixture models
192 Citations2003Fábio Gagliardi Cozman, Ira L. Cohen +1 more
This paper analyzes the performance of semi-supervised learning of mixture models and shows that unlabeled data can lead to an increase in classification error even in situations where additional labeled data would decrease classification error.
A continuation method for semi-supervised SVMs
185 Citations2006Olivier Chapelle, Mingmin Chi +1 more
This paper proposes to use a global optimization technique known as continuation to alleviate the problem of non-convex optimization of S3VMs, which often results in suboptimal performances.
Lecture notes in computer scienceMulti-label Image Segmentation for Medical Applications Based on Graph-Theoretic Electrical Potentials
184 Citations2004Leo Grady, Gareth Funka-Lea
A novel method is proposed for performing multi-label, semi-automated image segmentation using combinatorial analogues of standard operators and principles from continuous potential theory, allowing it to be applied in arbitrary dimension.
Learning & BehaviorBayesian approaches to associative learning: From passive to active learning
179 Citations2008John K. Kruschke
The first part of this article reviews two Bayesian accounts of backward blocking, a phenomenon that is challenging for many traditional theories and focuses on two formalizations of optimal active learning: maximizing either the expected information gain or the probability gain.
Érudit documents and data repository (Érudit Consortium, University of Montreal)Efficient Non-Parametric Function Induction in Semi-Supervised Learning
173 Citations2004Yoshua Bengio, Olivier Delalleau +1 more
Experiments show that the proposed non-parametric algorithms which provide an estimated continuous label for the given unlabeled examples are extended to function induction algorithms that correspond to the minimization of a regularization criterion applied to an out-of-sample example, and happens to have the form of a Parzen windows regressor.
Harmonic mixtures
169 Citations2005Xiaojin Zhu, John Lafferty
Experimental results show that this approach preserves the accuracy of purely graph-based transductive methods when the data has "manifold structure," and at the same time achieves inductive learning with significantly reduced computational cost.
On the relation between multi-instance learning and semi-supervised learning
167 Citations2007Zhi‐Hua Zhou, Junming Xu
The MissSVM algorithm is proposed which addresses multi- instance learning using a special semi-supervised support vector machine and is competitive with state-of-the-art multi-instance learning algorithms.
KiltHub RepositoryNonparametric Transforms of Graph Kernels for Semi-Supervised Learning
164 Citations2018Xiaojin Zhu, Jaz Kandola +2 more
An algorithm based on convex optimization for constructing kernels for semi-supervised learning that incorporates order constraints during optimization results in flexible kernels and avoids the need to choose among different parametric forms.
Simple, robust, scalable semi-supervised learning via expectation regularization
162 Citations2007Gideon Mann, Andrew McCallum
Expectation regularization is presented, a semi-supervised learning method for exponential family parametric models that augments the traditional conditional label-likelihood objective function with an additional term that encourages model predictions on unlabeled data to match certain expectations---such as label priors.
Efficient co-regularised least squares regression
154 Citations2006Ulf Brefeld, Thomas Gärtner +2 more
This paper investigates a semi-supervised least squares regression algorithm based on the co-learning approach that shows a significant error reduction by co-regularisation and a large runtime improvement for the semi-parametric approximation.
ACM Transactions on Information SystemsEnhancing relevance feedback in image retrieval using unlabeled data
154 Citations2006Zhi‐Hua Zhou, Kejia Chen +1 more
Experiments show that using semisupervised learning and active learning simultaneously in CBIR is beneficial, and the proposed method achieves better performance than some existing methods.
Unsupervised and semi-supervised multi-class support vector machines
151 Citations2005Linli Xu, Dale Schuurmans
A principled approach to unsupervised SVM training is presented by formulating convex relaxations of the natural training criterion: find a labeling that would yield an optimal SVM classifier on the resulting training data.
The Role of Unlabeled Data in Supervised Learning
150 Citations2004Tom M. Mitchell
It is argued that models of human and animal learning should consider more strongly the potential role of unlabeled data, and that many natural learning problems fit the problem class identified in this paper.
Neural Information Processing SystemsSemi-supervised Learning via Gaussian Processes
149 Citations2004Neil D. Lawrence, Michael I. Jordan
A probabilistic approach to learning a Gaussian Process classifier in the presence of unlabeled data using a "null category noise model" (NCNM) inspired by ordered categorical noise models.
The MIT Press eBooksManifold Denoising
148 Citations2007Matthias Hein, Markus Maier
The presented denoising algorithm is based on a graph-based diffusion process of the point sample and is analyzed using recent results about the convergence of graph Laplacians to improve the results of a semi-supervised learning algorithm.
Semisupervised Learning for Computational Linguistics
146 Citations2007Steven Abney
Graph transduction via alternating minimization
136 Citations2008Jun Wang, Tony Jebara +1 more
This paper introduces a propagation algorithm that more reliably minimizes a cost function over both a function on the graph and a binary label matrix and achieves substantial improvement in accuracy compared to state of the art semi-supervised methods.
…
