login

Web-Data Augmented Language Models for Mandarin Conversational Speech Recognition

Published 11 October 2006
T. Ng, Mari Ostendorf, Mei-Yuh Hwang, Man-Hung Siu, Ivan Bulyko, Xin Lei
Citations49

TL;DR

Experiments in recognizing Mandarin telephone conversations show that the use of filtered Web data leads to a 28% reduction in perplexity and 7% reductionIn character error rate, with most of the gain due to the general filtered WebData.

Abstract

Lack of data is a problem in training language models for conversational speech recognition, particularly for languages other than English. Experiments in English have successfully used web-based text collection targeted for a conversational style to augment small sets of transcribed speech; here we look at extending these techniques to Mandarin. In addition, we investigate different techniques for topic adaptation. Experiments in recognizing Mandarin telephone conversations show that use of filtered web data leads to a 28% reduction in perplexity and 7% reduction in character error rate, with most of the gain due to the general filtered web data.

Keywords

Computer Science