Introduction to Semi-Supervised Learning
Synthesis lectures on artificial intelligence and machine learningPublished 1 January 2009
Xiaojin Zhu, Andrew B. Goldberg
Citations1,024
SJR quartileQ4
SJR score0.23
SNIP2.20
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
Abstract
Semi-supervised learning is a learning paradigm concerned with the study of how computers and natural systems such as humans learn in the presence of both labeled and unlabeled data. Traditionally, le
Keywords
Computer Science
TechnometricsStatistical Learning Theory
26,913 Citations1999Yuhai Wu, Vladimir Vapnik
Presenting a method for determining the necessary and sufficient conditions for consistency of learning process, the author covers function estimates from small data pools, applying these estimations to real-life problems, and much more.
Springer series in statisticsThe Elements of Statistical Learning
24,344 Citations2001Trevor Hastie, J. Friedman +1 more
Choice Reviews OnlineArtificial intelligence: a modern approach
22,205 Citations1995Stuart Russell, Peter Norvig +2 more
Cambridge University Press eBooksAn Introduction to Support Vector Machines and Other Kernel-based Learning Methods
13,883 Citations2000Nello Cristianini, John Shawe‐Taylor
Mathematical ProgrammingOn the limited memory BFGS method for large scale optimization
8,529 Citations1989Dong C. Liu, Jorge Nocedal
The numerical tests indicate that the L-BFGS method is faster than the method of Buckley and LeNir, and is better able to use additional storage to accelerate convergence, and the convergence properties are studied to prove global convergence on uniformly convex problems.
Cambridge University Press eBooksKernel Methods for Pattern Analysis
6,598 Citations2004John Shawe‐Taylor, Nello Cristianini
This book provides an easy introduction for students and researchers to the growing field of kernel-based pattern analysis, demonstrating with examples how to handcraft an algorithm or a kernel for a new specific application, and covering all the necessary conceptual and mathematical tools to do so.
International Conference on Neural Information ProcessingAdvances in kernel methods: support vector learning
5,815 Citations1999Bernhard Schölkopf, Christopher J. C. Burges +1 more
Support vector machines for dynamic reconstruction of a chaotic system, Klaus-Robert Muller et al pairwise classification and support vector machines, Ulrich Kressel.
Combining labeled and unlabeled data with co-training
5,604 Citations1998Avrim Blum, Tom M. Mitchell
The MIT Press eBooksSemi-Supervised Learning
4,308 Citations2006Olivier Chapelle, Schölkopf, B. +1 more
This first comprehensive overview of semi-supervised learning presents state-of-the-art algorithms, a taxonomy of the field, selected applications, benchmark experiments, and perspectives on ongoing and future research.
A sentimental education
3,343 Citations2004Bo Pang, Lillian Lee
A novel machine-learning method is proposed that applies text-categorization techniques to just the subjective portions of the document, which greatly facilitates incorporation of cross-sentence contextual constraints.
Communications of the ACMA theory of the learnable
3,268 Citations1984Leslie G. Valiant
This paper regards learning as the phenomenon of knowledge acquisition in the absence of explicit programming, and gives a precise methodology for studying this phenomenon from a computational viewpoint.
A comparison of event models for naive bayes text classification
3,224 Citations1998Andrew McCallum, Kamal Nigam
It is found that the multi-variate Bernoulli performs well with small vocabulary sizes, but that the multinomial performs usually performs even better at larger vocabulary sizes--providing on average a 27% reduction in error over the multi -variateBernoulli model at any vocabulary size.
Machine LearningText Classification from Labeled and Unlabeled Documents using EM
2,749 Citations2000Kamal Nigam, Andrew Kachites McCallum +2 more
This paper shows that the accuracy of learned text classifiers can be improved by augmenting a small number of labeled training documents with a large pool of unlabeled documents, and presents two extensions to the algorithm that improve classification accuracy under these conditions.
Labeling images with a computer game
2,222 Citations2004Luis von Ahn, Laura Dabbish
A new interactive system: a game that is fun and can be used to create valuable output that addresses the image-labeling problem and encourages people to do the work by taking advantage of their desire to be entertained.
SWITCHBOARD: telephone speech corpus for research and development
2,140 Citations1992J. Godfrey, E. Holliman +1 more
The MIT Press eBooksAn Introduction to Computational Learning Theory
1,733 Citations1994Michael Kearns, Umesh Vazirani
The probably approximately correct learning model Occam's razor the Vapnik-Chervonenkis dimension weak and strong learning learning in the presence of noise inherent unpredictability reducibility in PAC learning learning finite automata is described.
Self-taught learning
1,533 Citations2007Rajat Raina, Alexis Battle +3 more
An approach to self-taught learning that uses sparse coding to construct higher-level features using the unlabeled data to form a succinct input representation and significantly improve classification performance.
ACM Transactions on GraphicsColorization using optimization
1,439 Citations2004Anat Levin, Dani Lischinski +1 more
The MIT Press eBooksLearning with Hypergraphs: Clustering, Classification, and Embedding
1,240 Citations2007Dengyong Zhou, Jiayuan Huang +1 more
This paper generalizes the powerful methodology of spectral clustering which originally operates on undirected graphs to hypergraphs, and further develop algorithms for hypergraph embedding and transductive classification on the basis of the spectral hypergraph clustering approach.
IEEE Transactions on Knowledge and Data EngineeringTri-training: exploiting unlabeled data using three classifiers
1,185 Citations2005Zhi‐Hua Zhou, Ming Li
Experiments on UCI data sets and application to the Web page classification task indicate that tri-training can effectively exploit unlabeled data to enhance the learning performance.
Springer texts in statisticsAll of Statistics
941 Citations2004Larry Wasserman
This book should certainly jump-start the process of using JMP effectively, and the authors have two other excellent users’ books for SAS that would provide an excellent basis for a JMP book similar to this one.
Lecture notes in computer scienceKernels and Regularization on Graphs
839 Citations2003Alexander J. Smola, Risi Kondor
It is shown that the class of positive, monotonically decreasing functions on the unit interval leads to kernels and corresponding regularization operators and can be found as a special case of the reasoning.
Diffusion Kernels on Graphs and Other Discrete Input Spaces
732 Citations2002Risi Kondor, John Lafferty
This paper proposes a general method of constructing natural families of kernels over discrete structures, based on the matrix exponentiation idea, and focuses on generating kernels on graphs, for which a special class of exponential kernels called diffusion kernels are proposed.
IEEE Transactions on Geoscience and Remote SensingThe effect of unlabeled samples in reducing the small sample size problem and mitigating the Hughes phenomenon
572 Citations1994B.M. Shahshahani, D. A. Landgrebe
By using additional unlabeled samples that are available at no extra cost, the performance may be improved, and therefore the Hughes phenomenon can be mitigated and therefore more representative estimates can be obtained.
ScienceThe Learning of Categories: Parallel Brain Systems for Item Memory and Category Knowledge
487 Citations1993Barbara J. Knowlton, Larry R. Squire
Amnesic patients and control subjects performed similarly at classifying novel patterns according to whether they belonged to the same category as a set of training patterns, and the amnesics were impaired at recognizing which dot patterns had been presented for training.
Learning from labeled and unlabeled data on a directed graph
403 Citations2005Dengyong Zhou, Jiayuan Huang +1 more
A general framework for learning from labeled and unlabeled data on a directed graph in which the structure of the graph including the directionality of the edges is considered, which generalizes the spectral clustering approach for undirected graphs.
Deep learning via semi-supervised embedding
387 Citations2008Jason Weston, Frédéric Ratle +1 more
It is shown how nonlinear embedding algorithms popular for use with shallow semi-supervised learning techniques such as kernel methods can be applied to deep multilayer architectures, either as a regularizer at the output layer, or on each layer of the architecture.
Trading convexity for scalability
361 Citations2006Ronan Collobert, Fabian H. Sinz +2 more
It is shown how concave-convex programming can be applied to produce faster SVMs where training errors are no longer support vectors, and much faster Transductive SVMs.
The MIT Press eBooksPAC Generalization Bounds for Co-training
287 Citations2002Sanjoy Dasgupta, Michael L. Littman +1 more
A new PAC-style bound on generalization error is given which justifies both the use of confidences — partial rules and partial labeling of the unlabeled data — and theUse of an agreement-based objective function as suggested by Collins and Singer.
An RKHS for multi-view learning and manifold co-regularization
224 Citations2008Vikas Sindhwani, David S. Rosenberg
This paper constructs a single RKHS with a data-dependent "co-regularization" norm that reduces these approaches to standard supervised learning and proposes a co- regularization based algorithmic alternative to manifold regularization that leads to major empirical improvements on semi-supervised tasks.
Semi-supervised learning of compact document representations with deep networks
223 Citations2008Marc ' Aurelio Ranzato, Martin Szummer
An algorithm to learn text document representations based on semi-supervised autoencoders that are stacked to form a deep network that can be trained efficiently on partially labeled corpora, producing very compact representations of documents, while retaining as much class information and joint word statistics as possible.
Perception & PsychophysicsOn the dominance of unidimensional rules in unsupervised categorization
203 Citations1999F. Gregory Ashby, Sarah Queller +1 more
In several experiments, observers tried to categorize stimuli constructed from two separable stimulus dimensions in the absence of any trial-by-trial feedback to support the hypothesis that people are constrained to use unidimensional rules.
Pattern Recognition LettersOn the exponential value of labeled samples
201 Citations1995Vittorio Castelli, Thomas M. Cover
The first labeled sample reduces the risk from 1 2 to 2R ∗ (1−R∗ ) and subsequent labeled samples in the training set reduce the probability of error exponentially fast to the Bayes risk.
A continuation method for semi-supervised SVMs
185 Citations2006Olivier Chapelle, Mingmin Chi +1 more
This paper proposes to use a global optimization technique known as continuation to alleviate the problem of non-convex optimization of S3VMs, which often results in suboptimal performances.
Lecture notes in computer scienceMulti-label Image Segmentation for Medical Applications Based on Graph-Theoretic Electrical Potentials
184 Citations2004Leo Grady, Gareth Funka-Lea
A novel method is proposed for performing multi-label, semi-automated image segmentation using combinatorial analogues of standard operators and principles from continuous potential theory, allowing it to be applied in arbitrary dimension.
Learning & BehaviorBayesian approaches to associative learning: From passive to active learning
179 Citations2008John K. Kruschke
The first part of this article reviews two Bayesian accounts of backward blocking, a phenomenon that is challenging for many traditional theories and focuses on two formalizations of optimal active learning: maximizing either the expected information gain or the probability gain.
Efficient co-regularised least squares regression
154 Citations2006Ulf Brefeld, Thomas Gärtner +2 more
This paper investigates a semi-supervised least squares regression algorithm based on the co-learning approach that shows a significant error reduction by co-regularisation and a large runtime improvement for the semi-parametric approximation.
ACM Transactions on Information SystemsEnhancing relevance feedback in image retrieval using unlabeled data
154 Citations2006Zhi‐Hua Zhou, Kejia Chen +1 more
Experiments show that using semisupervised learning and active learning simultaneously in CBIR is beneficial, and the proposed method achieves better performance than some existing methods.
The MIT Press eBooksManifold Denoising
148 Citations2007Matthias Hein, Markus Maier
The presented denoising algorithm is based on a graph-based diffusion process of the point sample and is analyzed using recent results about the convergence of graph Laplacians to improve the results of a semi-supervised learning algorithm.
Semisupervised Learning for Computational Linguistics
146 Citations2007Steven Abney
Graph transduction via alternating minimization
136 Citations2008Jun Wang, Tony Jebara +1 more
This paper introduces a propagation algorithm that more reliably minimizes a cost function over both a function on the graph and a binary label matrix and achieves substantial improvement in accuracy compared to state of the art semi-supervised methods.
Semi-supervised learning with very few labeled training examples
135 Citations2007Zhi‐Hua Zhou, De‐Chuan Zhan +1 more
By taking advantages of the correlations between the views using canonical component analysis, the proposed method can perform semi-supervised learning with only one labeled training example.
The MIT Press eBooksBranch and Bound for Semi-Supervised Support Vector Machines
129 Citations2007Olivier Chapelle, Vikas Sindhwani +1 more
Empirical evidence suggests that the globally optimal solution of S3VMs modulo local minima problems in current implementations can return excellent generalization performance in situations where other implementations fail completely.
The MIT Press eBooksOn Transductive Regression
113 Citations2007Corinna Cortes, Mehryar Mohri
This paper presents explicit VC-dimension error bounds for transductive regression that hold for all bounded loss functions and coincide with the tight classification bounds of Vapnik when applied to classification.
Word sense disambiguation using label propagation based semi-supervised learning
111 Citations2005Zheng-Yu Niu, Donghong Ji +1 more
This paper investigates a label propagation based semi-supervised learning algorithm for WSD, which combines labeled and unlabeled data in learning process to fully realize a global consistency assumption: similar examples should have similar labels.
Maximum Margin Semi-Supervised Learning for Structured Variables
99 Citations2005Yasemin Altün, David McAllester +1 more
This paper presents a discriminative approach that utilizes the intrinsic geometry of input patterns revealed by unlabeled data points and derives a maximum-margin formulation of semi-supervised learning for structured variables.
Journal of Artificial Intelligence ResearchLearning From Labeled And Unlabeled Data: An Empirical Study Across Techniques And Domains
97 Citations2005N. V. Chawla, G. Karakoulas
This paper presents an empirical study of various semi-supervised learning techniques on a variety of datasets, and attempts to answer various questions such as the effect of independence or relevance amongst features, theeffect of the size of the labeled and unlabeled sets and the effects of noise.
Statistical machine translation with word- and sentence-aligned parallel corpora
92 Citations2004Chris Callison-Burch, David Talbot +1 more
It is shown that significant improvements in the alignment and translation quality of such models can be achieved by additionally including word-aligned data during training by incorporating word-level alignments into the parameter estimation of the IBM models.
Probabilistic Modeling for Face Orientation Discrimination: Learning from Labeled and Unlabeled Data
87 Citations1998Shumeet Baluja
Semi-supervised learning for structured output variables
80 Citations2006Ulf Brefeld, Tobias Scheffer
The co-training approach is based on the principle of maximizing the consensus among multiple independent hypotheses and developed into a semi-supervised support vector learning algorithm for joint input output spaces and arbitrary loss functions.
Springer eBooksKernel Based Algorithms for Mining Huge Data Sets: Supervised, Semi-supervised, and Unsupervised Learning (Studies in Computational Intelligence)
74 Citations2006Te-Ming Huang, Vojislav Kecman +1 more
Content analysis of the data revealed that 25% of the women did not feel physically recovered from childbirth at 6 months postpartum, and pregnant women need more information about lifestyle adjustments after childbirth.
IEEE Transactions on Pattern Analysis and Machine IntelligenceSemisupervised Learning for a Hybrid Generative/Discriminative Classifier based on the Maximum Entropy Principle
73 Citations2008Akinori Fujino, N. Ueda +1 more
The experimental results for four text data sets confirmed that the generalization ability of the hybrid classifier was much improved by using a large number of unlabeled samples for training when there were too few labeled samples to obtain good performance.
On multi-view active learning and the combination with semi-supervised learning
73 Citations2008Wei Wang, Zhi‐Hua Zhou
The empirical behavior of the two paradigms is studied, which verifies that the combination of multi-view active learning and semi-supervised learning is efficient.
Machine LearningJoint feature re-extraction and classification using an iterative semi-supervised support vector machine algorithm
70 Citations2007Yuanqing Li, Cuntai Guan
An iterative semi-supervised SVM algorithm embedded with feature re-extraction that can be used to extract three features reliably and perform classification simultaneously in cases where the training data set is small.
The MIT Press eBooksLarge-Scale Sparsified Manifold Regularization
70 Citations2007Ivor W. Tsang, James T. Kwok
This paper integrates manifold regularization with the core vector machine and produces sparse solutions with low time and space complexities by using a sparsified manifold regularizer and formulating as a center-constrained minimum enclosing ball problem.
IEEE Transactions on Pattern Analysis and Machine IntelligenceQuery by Transduction
57 Citations2008Shen-Shyang Ho, Harry Wechsler
This paper proposes Query-by-Transduction (QBT) as a novel active learning algorithm that compares favorably, in terms of mean generalization, against random sampling, committee-based active learning, margin-basedactive learning, and QBC in the stream-based setting.
The MIT Press eBooksHyperparameter Learning for Graph Based Semi-supervised Learning Algorithms
57 Citations2007Xinhua Zhang, Wee Sun Lee
This paper proposes a graph learning method for the harmonic energy minimization method by minimizing the leave-one-out prediction error on labeled data points and designed an efficient algorithm which significantly accelerates the calculation of the gradient by applying the matrix inversion lemma and using careful pre-computation.
Lecture notes in computer scienceGeneralization Error Bounds Using Unlabeled Data
55 Citations2005Matti Kääriäinen
Two new methods for obtaining generalization error bounds in a semi-supervised setting based on approximating the disagreement probability of pairs of classifiers using unlabeled data are presented.
Stability of transductive regression algorithms
49 Citations2008Corinna Cortes, Mehryar Mohri +2 more
The notion of algorithmic stability is used to derive novel generalization bounds for several families of transductive regression algorithms, both by using convexity and closed-form solutions.
Lecture notes in computer scienceMulti-view Discriminative Sequential Learning
47 Citations2005Ulf Brefeld, Christoph Büscher +1 more
The multi-view approach is based on the principle of maximizing the consensus among multiple independent hypotheses and develops this principle into a semi-supervised hidden Markov perceptron, and the resulting procedures utilize unlabeled data effectively and discriminate more accurately than their purely supervised counterparts.
The asymptotics of semi-supervised learning in discriminative probabilistic models
46 Citations2008Nataliya Sokolovska, Olivier Cappé +1 more
An original methodology for using unlabeled data through the design of a simple semi-supervised objective function is introduced and it is proved that the corresponding semi- supervised estimator is asymptotically optimal.
Machine LearningMetric-Based Methods for Adaptive Model Selection and Regularization
46 Citations2002Dale Schuurmans, Finnegan Southey
A general approach to model selection and regularization that exploits unlabeled data to adaptively control hypothesis complexity in supervised learning tasks and derives a general training criterion for supervised learning that adjusts its regularization level to the specific set of training data received, and performs well on a variety of regression and conditional density estimation tasks.
Large scale manifold transduction
45 Citations2008Michael Karlen, Jason Weston +2 more
This work shows how the regularizer of Transductive Support Vector Machines can be trained by stochastic gradient descent for linear models and multi-layer architectures, and proposes a natural generalization of the TSVM loss function that takes into account neighborhood and manifold information directly.
Memory & CognitionA high-distortion enhancement effect in the prototype-learning paradigm: Dramatic effects of category learning during test
43 Citations2007Safa R. Zaki, Robert M. Nosofsky
The results provide dramatic evidence of the role of learning during transfer in this task and force a reevaluation of the dominant current interpretation of the steep typicality gradient.
Attention Perception & PsychophysicsSemisupervised category learning: The impact of feedback in learning the information-integration task
29 Citations2009Katleen Vandist, Maarten De Schryver +1 more
This article uses a semisupervised classification paradigm, in which feedback is given after a prespecified percentage of trials only, and shows that in both the 100% and 50% conditions, participants were able to achieve maximum accuracy.
arXiv (Cornell University)Analysis of Semi-Supervised Learning with the Yarowsky Algorithm
23 Citations2012Gholam Reza Haffari, Anoop Sarkar
Machine LearningLarge margin vs. large volume in transductive learning
20 Citations2008Ran El‐Yaniv, Dmitry Pechyony +1 more
ManifoldBoost
18 Citations2008Nicolas Loeff, David Forsyth +1 more
A manifold learning framewor that naturally accommodates supervised learning, partially supervised learning and unsupervised clustering as particular cases is described and the performance is at the state of the art on many standard semi-supervised learning benchmarks.
The MIT Press eBooksProbabilistic Semi-Supervised Clustering with Constraints
16 Citations2006Basu Sugato, Bilenko Mikhail +2 more
Pattern Recognition LettersEffective transductive learning via objective model selection
15 Citations2005Ran El‐Yaniv, Leonid Gerzon
Empirical examination of a recent transductive learning approach based on clustering, implemented with 'spectral clustering', on a suite of benchmark datasets from the UCI repository indicates that the new approach is effective and comparable with one of the best known transductive learning algorithms to-date.
Transactions of the Japanese Society for Artificial IntelligenceA Hybrid Generative/Discriminative Classifier Design for Semi-supervised Learing
9 Citations2006Akinori Fujino, Naonori Ueda +1 more
A hybrid classifier is constructed by combining both the generative and bias correction models based on the maximum entropy principle, where the combination weights of these models are determined so that the class labels of labeled samples are as correctly predicted as possible.
The MIT Press eBooksAn Augmented PAC Model for Semi-Supervised Learning
7 Citations2006Balcan Maria-Florina, Blum Avrim
The MIT Press eBooksSemi-Supervised Learning with Conditional Harmonic Mixing
5 Citations2006Burges Christopher J. C., Platt John C.
This chapter introduces a general probabilistic formulation called `Conditional Harmonic Mixing’, in which the links are directed, a conditional probability matrix is associated with each link, and where the numbers of classes can vary from node to node.
…
