login

SRI's 2004 NIST Speaker Recognition Evaluation System

Published 11 October 2006
Sachin Kajarekar, Luciana Ferrer, E. Shriberg, Kemal Sönmez, Andreas Stolcke, Anand Venkataraman
Citations50

TL;DR

A system that uses discriminant features from cepstral coefficients, and systems that use discriminant models from word n-grams and syllable-based NERF n- grams together with a cEPstral baseline system are evaluated.

Abstract

The paper describes our recent efforts in exploring longer-range features and their statistical modeling techniques for speaker recognition. In particular, we describe a system that uses discriminant features from cepstral coefficients, and systems that use discriminant models from word n-grams and syllable-based NERF n-grams. These systems together with a cepstral baseline system are evaluated on the 2004 NIST speaker recognition evaluation dataset. The effect of the development set is measured using two different datasets, one from Switchboard databases and another from the FISHER database. Results show that the difference between the development and evaluation sets affects the performance of the systems only when more training data is available. Results also show that systems using longer-range features combined with the baseline result in about a 31% improvement with 1-side training over the baseline system and about a 61% improvement with 8-side training over the baseline system.

Keywords

Computer Science