Cross-Language Information Retrieval: A System for Comparable Corpus Querying
Generate an AI Snapshot to get a quick, structured summary of this paper.
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
A system that has been designed to process and query comparable text corpora, i.e. collections of texts from pairs or multiples of languages referring to the same domain, which could have applications in the fields of terminology and crosslingual document retrieval.
Abstract
We describe a system that has been designed to process and query comparable text corpora, i.e. collections of texts from pairs or multiples of languages referring to the same domain. The first version of the system has been developed to retrieve natural language lexical equivalents from sets of sublanguage texts in English and Italian; given the necessary lexical and morphological components it could be extended to cover other languages. The initial implementation was made with the needs of language scholars in mind; however, the system could have applications in the fields of terminology and crosslingual document retrieval.
