login

Visual Information Retrieval

SpringerReferencePublished 20 January 2012
Alberto Del Bimbo, Marko Grobelnik, Peter Jackson, Thomson Corporation
Citations740

TL;DR

This work applied the Kernel Canonical Correlation Analysis to the Japanese-English cross-language information retrieval and classification and the results were encouraging.

Abstract

We present a new, ontology-based approach to the automatic text categorization.An important and novel aspect of this approach is that our categorization method does not require a training set, which is in contrast to the traditional statistical and probabilistic methods that require a set of preclassified documents in order to train the classifier.In our approach, the ontology, which holds the schema, including the domain entities organized into categories and interconnected by relationships, as well as instances and linkages among them, effectively becomes the classifier for the categories of the domain concepts.After a document is converted into a thematic graph of entities, the ontological classification of the entities in the graph is then analyzed in order to determine the overall categorization of the thematic graph, and as a result, of the document.In presented experiments, we used an RDF ontology constructed from the full English version of Wikipedia, a Web-based encyclopedia.The experiments, conducted on a collection of news articles, show that our training-less categorization method has achieved a satisfactory overall accuracy, in one experiment nearly identical to a selected traditional categorization method.

Keywords

Computer Science