Visual Information Retrieval
Generate an AI Snapshot to get a quick, structured summary of this paper.
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
This work applied the Kernel Canonical Correlation Analysis to the Japanese-English cross-language information retrieval and classification and the results were encouraging.
Abstract
We present a new, ontology-based approach to the automatic text categorization.An important and novel aspect of this approach is that our categorization method does not require a training set, which is in contrast to the traditional statistical and probabilistic methods that require a set of preclassified documents in order to train the classifier.In our approach, the ontology, which holds the schema, including the domain entities organized into categories and interconnected by relationships, as well as instances and linkages among them, effectively becomes the classifier for the categories of the domain concepts.After a document is converted into a thematic graph of entities, the ontological classification of the entities in the graph is then analyzed in order to determine the overall categorization of the thematic graph, and as a result, of the document.In presented experiments, we used an RDF ontology constructed from the full English version of Wikipedia, a Web-based encyclopedia.The experiments, conducted on a collection of news articles, show that our training-less categorization method has achieved a satisfactory overall accuracy, in one experiment nearly identical to a selected traditional categorization method.
