login

Cross-Language Information Retrieval: A System for Comparable Corpus Querying

Published 1 January 1998
Eugenio Picchi, C Peters
Citations51

TL;DR

A system that has been designed to process and query comparable text corpora, i.e. collections of texts from pairs or multiples of languages referring to the same domain, which could have applications in the fields of terminology and crosslingual document retrieval.

Abstract

We describe a system that has been designed to process and query comparable text corpora, i.e. collections of texts from pairs or multiples of languages referring to the same domain. The first version of the system has been developed to retrieve natural language lexical equivalents from sets of sublanguage texts in English and Italian; given the necessary lexical and morphological components it could be extended to cover other languages. The initial implementation was made with the needs of language scholars in mind; however, the system could have applications in the fields of terminology and crosslingual document retrieval.

Keywords

Computer Science