login

Cross-language information retrieval based on parallel texts and automatic mining of parallel texts from the Web

Published 1 August 1999
Jian‐Yun Nie, Michel Simard, Pierre Isabelle, Richard Durand
Citations310

TL;DR

It is shown that using a probabilistic model, it is able to obtain performances close to those using an MT system, and the possibility of automatically gather parallel texts from the Web in an attempt to construct a reasonable training corpus is investigated.

Abstract

This paper describes the use of a probabilistic translation model to cross-language IR (CLIR). The performance of this approach is compared with that using machine translation (MT). It is shown that using a probabilistic model, we are able to obtain performances close to those using an MT system. In addition, we also investigated the possibility of automatically gather parallel texts from the Web in an attempt to construct a reasonable training corpus. The result is very encouraging. We showed that in several tests, such a training corpus is as good as a manually constructed one for CLIR purposes.

Keywords

Computer Science