TagProp: Discriminative metric learning in nearest neighbor models for image auto-annotation
Published 1 September 2009Open access
Matthieu Guillaumin, Thomas Mensink, Jakob Verbeek, Cordelia Schmid
Citations698
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
This work proposes TagProp, a discriminatively trained nearest neighbor model that allows the integration of metric learning by directly maximizing the log-likelihood of the tag predictions in the training set, and introduces a word specific sigmoidal modulation of the weighted neighbor tag predictions to boost the recall of rare words.
Abstract
International audience
Keywords
Computer Science
Beyond Bags of Features: Spatial Pyramid Matching for Recognizing Natural Scene Categories
7,929 Citations2006Svetlana Lazebnik, C. Schmid +1 more
This paper presents a method for recognizing scene categories based on approximate global geometric correspondence that exceeds the state of the art on the Caltech-101 database and achieves high accuracy on a large database of fifteen natural scene categories.
International Journal of Computer VisionModeling the Shape of the Scene: A Holistic Representation of the Spatial Envelope
6,410 Citations2001Aude Oliva, Antonio Torralba
The performance of the spatial envelope model shows that specific information about object shape or identity is not a requirement for scene categorization and that modeling a holistic representation of the scene informs about its probable semantic category.
Labeling images with a computer game
2,222 Citations2004Luis von Ahn, Laura Dabbish
A new interactive system: a game that is fun and can be used to create valuable output that addresses the image-labeling problem and encourages people to do the work by taking advantage of their desire to be entertained.
Lecture notes in computer scienceObject Recognition as Machine Translation: Learning a Lexicon for a Fixed Image Vocabulary
1,608 Citations2002Pınar Duygulu, Kobus Barnard +2 more
This work shows how to cluster words that individually are difficult to predict into clusters that can be predicted well, and cannot predict the distinction between train and locomotive using the current set of features, but can predict the underlying concept.
Automatic image annotation and retrieval using cross-media relevance models
1,146 Citations2003Jiwoon Jeon, Victor Lavrenko +1 more
The approach shows the usefulness of using formal information retrieval models for the task of image annotation and retrieval by assuming that regions in an image can be described using a small vocabulary of blobs.
SVM-KNN: Discriminative Nearest Neighbor Classification for Visual Category Recognition
1,116 Citations2006Hao Zhang, Alexander C. Berg +2 more
This work considers visual category recognition in the framework of measuring similarities, or equivalently perceptual distances, to prototype examples of categories and proposes a hybrid of these two methods which deals naturally with the multiclass setting, has reasonable computational complexity both in training and at run time, and yields excellent results in practice.
IEEE Transactions on Pattern Analysis and Machine IntelligenceSupervised Learning of Semantic Classes for Image Annotation and Retrieval
866 Citations2007Gustavo Carneiro, Antoni B. Chan +2 more
The supervised formulation is shown to achieve higher accuracy than various previously published methods at a fraction of their computational cost and to be fairly robust to parameter tuning.
Multiple Bernoulli relevance models for image and video annotation
813 Citations2004Siwei Feng, R. Manmatha +1 more
This work shows how it can do both automatic image annotation and retrieval (using one word queries) from images and videos using a multiple Bernoulli relevance model, which significantly outperforms previously reported results on the task of image and video annotation.
Is that you? Metric learning approaches for face identification
758 Citations2009Matthieu Guillaumin, Jakob Verbeek +1 more
Two methods for learning robust distance measures are presented: a logistic discriminant approach which learns the metric from a set of labelled image pairs (LDML) and a nearest neighbour approach which computes the probability for two images to belong to the same class (MkNN).
ScholarWorks@UMassAmherst (University of Massachusetts Amherst)A Model for Learning the Semantics of Pictures
681 Citations2003Victor Lavrenko, R. Manmatha +1 more
An approach to learning the semantics of images which allows us to automatically annotate an image with keywords and to retrieve images based on text queries using a formalism that models the generation of annotated images.
Metric Learning by Collapsing Classes
663 Citations2005Amir Globerson, Sam T. Roweis
An algorithm for learning a quadratic Gaussian metric (Mahalanobis distance) for use in classification tasks and discusses how the learned metric may be used to obtain a compact low dimensional feature representation of the original input space, allowing more efficient classification with very little reduction in performance.
IEEE Transactions on Pattern Analysis and Machine IntelligenceReal-Time Computerized Annotation of Pictures
511 Citations2008Jia Li, James Z. Wang
Automatic multimedia cross-modal correlation discovery
488 Citations2004Jia-Yu Pan, Hyung-Jeong Yang +2 more
A novel, graph-based approach, "MMG", to discover cross-modal correlations across the media in a collection of multimedia objects, where it outperforms domain specific, fine-tuned methods by up to 10 percentage points in captioning accuracy.
Matching Words and Pictures
459 Citations2003Kobus Barnard, Pinar Duygulu +8 more
Lecture notes in computer scienceA New Baseline for Image Annotation
444 Citations2008Ameesh Makadia, Vladimir Pavlović +1 more
This work introduces a new baseline technique for image annotation that treats annotation as a retrieval problem and outperforms the current state-of-the-art methods on two standard and one large Web dataset.
Lecture notes in computer scienceColoring Local Feature Extraction
432 Citations2006Joost van de Weijer, Cordelia Schmid
The results show that color descriptors remain reliable under photometric and geometrical changes, and with decreasing image quality, and for all experiments a combination of color and shape outperforms a pure shape-based approach.
Proceedings of the 26th annual international ACM SIGIR conference on Research and development in informaion retrieval - SIGIR '03Automatic image annotation and retrieval using cross-media relevance models
418 Citations2003Jiwoon Jeon, Victor Lavrenko +1 more
IEEE Transactions on Pattern Analysis and Machine IntelligenceA Discriminative Kernel-Based Approach to Rank Images from Text Queries
325 Citations2008David Grangier, Samy Bengio
This paper introduces a discriminative model for the retrieval of images from text queries that formalizes the retrieval task as a ranking problem, and introduces a learning procedure optimizing a criterion related to the ranking performance.
Proceedings of SPIE, the International Society for Optical Engineering/Proceedings of SPIE<title>Image annotation using SVM</title>
259 Citations2003Claudio Cusano, Gianluigi Ciocca +1 more
An innovative image annotation tool for classifying image regions in one of seven classes - sky, skin, vegetation, snow, water, ground, and buildings - or as unknown is described.
PLSA-based image auto-annotation
219 Citations2004Florent Monay, Daniel Gática-Pérez
A new way of modeling multi-modal co-occurrences is proposed, constraining the definition of the latent space to ensure its consistency in semantic terms (words), while retaining the ability to jointly model visual information.
Pattern RecognitionImage annotation via graph learning
197 Citations2008Jing Liu, Mingjing Li +3 more
A Nearest Spanning Chain (NSC) method is proposed to construct the image-based graph, whose edge-weights are derived from the chain-wise statistical information instead of the traditional pairwise similarities.
Lecture notes in computer scienceAutomated Image Annotation Using Global Features and Robust Nonparametric Density Estimation
138 Citations2005Alexei Yavlinsky, Edward Schofield +1 more
It is shown that under this framework quite simple image properties such as global colour and texture distributions provide a strong basis for reliably annotating images.
Learning distance functions for image retrieval
117 Citations2004Tomer Hertz, Aharon Bar-Hillel +1 more
The main contribution is a distance learning method, which combines boosting hypotheses over the product space with a weak learner based on partitioning the original feature space, which outperforms existing metric learning methods, which are based an learning a Mahalanobis distance.
Lecture notes in computer scienceAn Inference Network Approach to Image Retrieval
98 Citations2004Donald Metzler, R. Manmatha
A model based on the Inference Network framework from information retrieval that employs a powerful query language that allows structured query operators, term weighting, and the combination of text and images within a query is proposed.
Victoria University Research Repository (Victoria University)Analysis and evaluation of visual information systems performance
82 Citations2007Michael Grubinger
A deeper understanding of the complex conditions and constraints associated with visual information identification, the accurate capturing of user requirements, the appropriate specification and complexity of user queries, the execution of searches, and the reliability of performance indicators is enabled.
2009 IEEE Conference on Computer Vision and Pattern RecognitionLearning a distance metric from multi-instance multi-label data
79 Citations2009Rong Jin, Shijun Wang +1 more
An iterative algorithm is proposed for MIML distance metric learning that first estimates the association between instances in a bag and its assigned class labels, and learns a distance metric from the estimated association by a discriminative analysis.
Annotating images and image objects using a hierarchical dirichlet process model
52 Citations2008Oksana Yakhnenko, Vasant Honavar
A nonparametric Bayesian model which provides a generalization for multi-model latent Dirichlet allocation model (MoM-LDA) used for similar problems in the past and performs just as well as or better than the MoM- LDA model (regardless of the choice of the number of clusters) for predicting labels of objects in images containing multiple objects.
Coherent image annotation by learning semantic distance
51 Citations2008Tao Mei, Yong Wang +3 more
A novel approach to image annotation which simultaneously learns a semantic distance by capturing the prior annotation knowledge and propagates the annotation of an image as a whole entity is proposed.
