Learning with Unlabeled Data and Its Application to Image Retrieval
Lecture notes in computer sciencePublished 1 January 2006
Zhi‐Hua Zhou
Citations38
SJR quartileQ2
SJR score0.35
SNIP0.55
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
Abstract
In many practical machine learning or data mining applications, unlabeled training examples are readily available but labeled ones are fairly expensive to obtain because labeling the examples require human effort. So, learning with unlabeled data has attracted much attention during the past few years. This paper shows that how such techniques can be helpful in a difficult task, content-based image retrieval, for improving the retrieval performance by exploiting images existing in the database.
Keywords
Computer Science
TechnometricsStatistical Learning Theory
26,913 Citations1999Yuhai Wu, Vladimir Vapnik
Presenting a method for determining the necessary and sufficient conditions for consistency of learning process, the author covers function estimates from small data pools, applying these estimations to real-life problems, and much more.
NeurocomputingAdvances in neural information processing systems 7
22,300 Citations1997Krzysztof J. Cios, Mark E. Shields
IEEE Transactions on Pattern Analysis and Machine IntelligenceContent-based image retrieval at the end of the early years
6,023 Citations2000A.W.M. Smeulders, Marcel Worring +3 more
The working conditions of content-based retrieval: patterns of use, types of pictures, the role of semantics, and the sensory gap are discussed, as well as aspects of system engineering: databases, system architecture, and evaluation.
Combining labeled and unlabeled data with co-training
5,604 Citations1998Avrim Blum, Tom M. Mitchell
Minds at UW (University of Wisconsin)Semi-Supervised Learning Literature Survey
3,868 Citations2005Xiaojin Zhu
The study clearly indicates that the common practice of stripwise precommercial thinning is unjustified, and the justification of heavy 'chessboard' thinning (with pruning) depends on whether the potential reduction in rotation length and the improvement in wood quality outweigh the discounted costs of pre-commercial thinning and selection and pruning of crop trees.
Neural ComputationCanonical Correlation Analysis: An Overview with Application to Learning Methods
3,304 Citations2004David R. Hardoon, Sándor Szedmák +1 more
A general method using kernel canonical correlation analysis to learn a semantic representation to web images and their associated text and compares orthogonalization approaches against a standard cross-representation retrieval technique known as the generalized vector space model is presented.
Machine LearningText Classification from Labeled and Unlabeled Documents using EM
2,749 Citations2000Kamal Nigam, Andrew Kachites McCallum +2 more
This paper shows that the accuracy of learned text classifiers can be improved by augmenting a small number of labeled training documents with a large pool of unlabeled documents, and presents two extensions to the algorithm that improve classification accuracy under these conditions.
Transductive Inference for Text Classification using Support Vector Machines
2,717 Citations1999Thorsten Joachims
An analysis of why TSVMs are well suited for text classi(cid:12)cation is presented, and an algorithm for training TSVMs e(cid:14)-ciently, handling 10,000 examples and more is proposed.
IEEE Transactions on Circuits and Systems for Video TechnologyRelevance feedback: a power tool for interactive content-based image retrieval
1,764 Citations1998Yong Rui, Thomas S. Huang +2 more
A relevance feedback based interactive retrieval approach that effectively takes into account the subjectivity of human perception of visual content and the gap between high-level concepts and low-level features in CBIR.
Query by committee
1,657 Citations1992H. Sebastian Seung, Manfred Opper +1 more
It is suggested that asymptotically finite information gain may be an important characteristic of good query algorithms, in which a committee of students is trained on the same data set.
IEEE Transactions on Knowledge and Data EngineeringTri-training: exploiting unlabeled data using three classifiers
1,185 Citations2005Zhi‐Hua Zhou, Ming Li
Experiments on UCI data sets and application to the Web page classification task indicate that tri-training can effectively exploit unlabeled data to enhance the learning performance.
A Sequential Algorithm for Training Text Classifiers
1,158 Citations1994David Lewis, William A. Gale
An algorithm for sequential sampling during machine learning of statistical classifiers was developed and tested on a newswire text categorization task and reduced by as much as 500-fold the amount of training data that would have to be manually classified to achieve a given level of effectiveness.
Multimedia SystemsRelevance feedback in image retrieval: A comprehensive review
823 Citations2003Xiang Sean Zhou, Thomas S. Huang
The nature of the relevance feedback problem in a continuous representation space in the context of content-based image retrieval is analyzed and a list of critical issues to consider when designing a relevance feedback algorithm is compiled.
IEEE Transactions on Geoscience and Remote SensingThe effect of unlabeled samples in reducing the small sample size problem and mitigating the Hughes phenomenon
572 Citations1994B.M. Shahshahani, D. A. Landgrebe
By using additional unlabeled samples that are available at no extra cost, the performance may be improved, and therefore the Hughes phenomenon can be mitigated and therefore more representative estimates can be obtained.
Enhancing Supervised Learning with Unlabeled Data
459 Citations2000Sally A. Goldman, Yan Zhou
A new semi-supervised learning method called co-learning that is designed to use unlabeled data to enhance standard supervised learning algorithms to leverage off the fact that they have different representations of the hypotheses and are likely to detect different patterns in labeled data.
Neural Information Processing SystemsA Mixture of Experts Classifier with Learning Based on Both Labelled and Unlabelled Data
289 Citations1996David J. Miller, Hasan S. Uyar
A classifier structure and learning algorithm that make effective use of unlabelled data to improve performance and is a "mixture of experts" structure that is equivalent to the radial basis function (RBF) classifier, but unlike RBFs, is amenable to likelihood-based training.
International Joint Conference on Artificial IntelligenceSemi-supervised regression with co-training
244 Citations2005Zhi‐Hua Zhou, Ming Li
Experiments show that COREG can effectively exploit unlabeled data to improve regression estimates and is proposed as a co-training style semi-supervised regression algorithm.
Selective Sampling with Redundant Views
169 Citations2000Ion Muslea, Steven Minton +1 more
The most general algorithm in the cotesting family, naive co-testing, is analyzed, which can be used with virtually any type of learner and may also boost the classification accuracy.
ACM Transactions on Information SystemsEnhancing relevance feedback in image retrieval using unlabeled data
154 Citations2006Zhi‐Hua Zhou, Kejia Chen +1 more
Experiments show that using semisupervised learning and active learning simultaneously in CBIR is beneficial, and the proposed method achieves better performance than some existing methods.
Lecture notes in computer scienceExploiting Unlabeled Data in Content-Based Image Retrieval
64 Citations2004Zhi‐Hua Zhou, Kejia Chen +1 more
The SSAIR (Semi-Supervised Active Image Retrieval) approach, which attempts to exploit unlabeled data to improve the performance of content-based image retrieval (CBIR), is proposed and experiments show that semi-supervised learning and active learning mechanisms are both beneficial to CBIR.
