login

Transliteration of proper names in cross-lingual information retrieval

Published 1 January 2003Open access
Paola Virga, Sanjeev Khudanpur
Citations172
View PDF

TL;DR

The application of statistical machine translation techniques to "translate" the phonemic representation of an English name to a sequence of initials and finals, commonly used sub-word units of pronunciation for Chinese in support of cross-lingual speech and text processing applications.

Abstract

We address the problem of transliterating English names using Chinese orthography in support of cross-lingual speech and text processing applications. We demonstrate the application of statistical machine translation techniques to "translate" the phonemic representation of an English name, obtained by using an automatic text-to-speech system, to a sequence of initials and finals, commonly used sub-word units of pronunciation for Chinese. We then use another statistical translation model to map the initial/final sequence to Chinese characters. We also present an evaluation of this module in retrieval of Mandarin spoken documents from the TDT corpus using English text queries.

Keywords

Computer Science