Language modeling for bag-of-visual words image categorization
Published 7 July 2008
Pierre Tirilly, Vincent Claveau, Patrick Gros
Citations111
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
Two ways of improving image classification based on bag-of-words representation are proposed and new techniques to eliminate useless words are proposed, one based on geometric properties of the keypoints, the other on the use of probabilistic Latent Semantic Analysis (pLSA).
Abstract
International audience
Keywords
Computer Science
Encyclopedia of Statistics in Behavioral SciencePrincipal Component Analysis
14,549 Citations2005Ian T. Jolliffe
IEEE Transactions on Pattern Analysis and Machine IntelligenceA performance evaluation of local descriptors
6,692 Citations2005Krystian Mikolajczyk, C. Schmid
Video Google: a text retrieval approach to object matching in videos
6,414 Citations2003Sivic, Zisserman
An approach to object and scene retrieval which searches for and localizes all the occurrences of a user outlined object in a video, represented by a set of viewpoint invariant region descriptors so that recognition can proceed successfully despite changes in viewpoint, illumination and partial occlusion.
Visual categorization with bags of keypoints
4,147 Citations2004G. Csurka
Scalable Recognition with a Vocabulary Tree
3,612 Citations2006D. Nistér, Henrik Stewénius
A recognition scheme that scales efficiently to a large number of objects and allows a larger and more discriminatory vocabulary to be used efficiently is presented, which it is shown experimentally leads to a dramatic improvement in retrieval quality.
International Journal of Computer VisionA Comparison of Affine Region Detectors
2,921 Citations2005Krystian Mikolajczyk, Tinne Tuytelaars +6 more
A snapshot of the state of the art in affine covariant region detectors, and compares their performance on a set of test images under varying imaging conditions to establish a reference test set of images and performance software so that future detectors can be evaluated in the same framework.
ACM SIGIR ForumA Language Modeling Approach to Information Retrieval
2,532 Citations2017Jay Ponte, W. Bruce Croft
It will be shown that probabilistic methods can be used to predict topic changes in the context of the task of new event detection and provide further proof of concept for the use of language models for retrieval tasks.
Learning Generative Visual Models from Few Training Examples: An Incremental Bayesian Approach Tested on 101 Object Categories
2,482 Citations2005Li Fei-Fei, Rob Fergus +1 more
Machine LearningUnsupervised Learning by Probabilistic Latent Semantic Analysis
2,443 Citations2001Thomas Hofmann
This paper proposes to make use of a temperature controlled version of the Expectation Maximization algorithm for model fitting, which has shown excellent performance in practice, and results in a more principled approach with a solid foundation in statistical inference.
Computer Speech & LanguageAn empirical study of smoothing techniques for language modeling
2,077 Citations1999Stanley F. Chen, Joshua Goodman
International Journal of Computer VisionLocal Features and Kernels for Classification of Texture and Object Categories: A Comprehensive Study
1,931 Citations2006Jianguo Zhang, M. Marszałek +2 more
A large-scale evaluation of an approach that represents images as distributions of features extracted from a sparse set of keypoint locations and learns a Support Vector Machine classifier with kernels based on two effective measures for comparing distributions, the Earth Mover’s Distance and the χ2 distance.
A language modeling approach to information retrieval
1,723 Citations1998Jay Ponte, W. Bruce Croft
This work proposes an approach to retrieval based on probabilistic language modeling and integrates document indexing and document retrieval into a single model, which significantly outperforms standard tf.idf weighting on two different collections and query sets.
N-gram-based text categorization
1,503 Citations1994William B. Cavnar, John M. Trenkle
An N-gram-based approach to text categorization that is tolerant of textual errors is described, which worked very well for language classification and worked reasonably well for classifying articles from a number of different computer-oriented newsgroups according to subject.
Image Classification using Random Forests and Ferns
1,235 Citations2007Anna Bosch, Andrew Zisserman +1 more
It is shown that selecting the ROI adds about 5% to the performance and, together with the other improvements, the result is about a 10% improvement over the state of the art for Caltech-256.
Discovering objects and their location in images
980 Citations2005Josef Šivic, Bryan Russell +3 more
This work treats object categories as topics, so that an image containing instances of several categories is modeled as a mixture of topics, and develops a model developed in the statistical text literature: probabilistic latent semantic analysis (pLSA).
Evaluating bag-of-visual-words representations in scene classification
812 Citations2007Jun Yang, Yu–Gang Jiang +2 more
This study provides an empirical basis for designing visual-word representations that are likely to produce superior classification performance and applies techniques used in text categorization to generate image representations that differ in the dimension, selection, and weighting of visual words.
Lecture notes in computer scienceScene Classification Via pLSA
768 Citations2006Anna Bosch, Andrew Zisserman +1 more
The classification performance under changes in the visual vocabulary and number of latent topics learnt is investigated, and a novel vocabulary using colour SIFT descriptors is developed using probabilistic Latent Semantic Analysis.
An empirical study of smoothing techniques for language modeling
754 Citations1996Stanley F. Chen, Joshua Goodman
Statistical language modeling using the CMU-cambridge toolkit
564 Citations1997Philip Clarkson, Roni Rosenfeld
The conventional language modeling technology, as implemented in the toolkit, is outlined, and the extra e(cid:14)ciency and functionality that the new toolkit provides as compared to previous software for this task is described.
The MIT Press eBooksFast Discriminative Visual Codebooks using Randomized Clustering Forests
456 Citations2007Frank Moosmann, Bill Triggs +1 more
This work introduces Extremely Randomized Clustering Forests - ensembles of randomly created clustering trees - and shows that these provide more accurate results, much faster training and testing and good resistance to background clutter in several state-of-the-art image classification tasks.
Discovering object categories in image collections
442 Citations2005Josef Šivic, Bryan Russell +3 more
Given a set of images containing multiple object categories, this work seeks to discover those categories and their image locations without supervision using generative models from the statistical text literature: probabilistic Latent Semantic Analysis (pLSA), and Latent Dirichlet Allocation (LDA).
Discovery of Collocation Patterns: from Visual Words to Visual Phrases
238 Citations2007Junsong Yuan, Ying Wu +1 more
A fast and principled solution to the discovery of significant spatial co-occurrent patterns using frequent itemset mining; a pattern summarization method that deals with the compositional uncertainties in visual phrases; and a top-down refinement scheme of the visual word lexicon by feeding back discovered phrases to tune the similarity measure through metric learning.
ACM Transactions on Information SystemsTheory of keyblock-based image retrieval
91 Citations2002Lei Zhu, A. Rao +1 more
A keyblock-based approach to content-based image retrieval where each image is encoded as a set of one-dimensional index codes linked to the keyblocks in the codebook, analogous to considering a text document as a linear list of keywords.
Scene Classification Using Bag-of-Regions Representations
91 Citations2007Demir Gokalp, Selim Aksoy
A novel region selection algorithm is proposed that identifies region types that are frequently found in a particular class of Scenes but rarely exist in other classes, and also consistently occur together in the same class of scenes.
Effective and efficient object-based image retrieval using visual phrases
71 Citations2006Qingfang Zheng, Weiqiang Wang +1 more
This paper draws an analogy between image retrieval and text retrieval and proposes a visual phrase-based approach to retrieve images containing desired objects and devise methods on how to construct visual phrases from images and how to encode the visual phrase for indexing and retrieval.
Visual language modeling for image classification
65 Citations2007Lei Wu, Mingjing Li +3 more
A visual language modeling method for content-based image classification that transforms each image into a matrix of visual words, and assumes that each visual word is conditionally dependent on its neighbors, which can utilize the spatial correlation ofVisual words effectively in image classification.
Latent Mixture Vocabularies for Object Categorization
48 Citations2006Diane Larlus, Frédéric Jurie
Flexible spatial models for grouping local image features
40 Citations2004Gustavo Carneiro, Allan D. Jepson
This work uses semi-local spatial constraints which allow for a greater range of shape deformations and its functionality is demonstrated in an exemplar-based object recognition system that deals well with severe non-rigid deformations.
Automatic Image Annotation through Multi-Topic Text Categorization
37 Citations2006Sheng Gao, Dehong Wang +1 more
A new framework for automatic image annotation through multi-topic text categorization using topic classifiers trained from images with multiple associations, including spatial, syntactic, or semantic relationship between images and concepts is proposed.
Using Language Models for Text Classification
24 Citations2004Jing Bai, Jian‐Yun Nie
This approach is a natural extension of the traditional Naive Bayes classifier, in which the Laplace smoothing is replaced by some more sophisticated smoothing methods, and an additional factor of smoothing scale is introduced according to the amount of training data of the class to improve the classification performance.
Learning Structured Appearance Models from Captioned Images of Cluttered Scenes
20 Citations2007Michael Jamieson, Afsaneh Fazly +3 more
A connected graph appearance model where vertices represent local features and edges encode spatial relationships is described, which uses the repetition of feature neighborhoods across training images and a measure of correspondence with caption words to guide the search for meaningful feature configurations.
Creation de vocabulaires visuels efficaces pour la categorisation d’images
8 Citations2005Diane Larlus, Gyuri Dorkó +1 more
Dublin City University Open Access Institutional Repository (Dublin City University)Discrete language models for video retrieval
7 Citations2005Kieran McDonald
This thesis proposes to model colour, edge and texture histogrambased features directly with discrete language models and this approach is compatible with further traditional visual feature representations and provides a consistent, effective and relatively efficient model for video retrieval.
