Speaker recognition: a tutorial
Proceedings of the IEEEPublished 1 January 1997Open access
J.P. Campbell
Citations1,626
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
A tutorial on the design and development of automatic speaker-recognition systems is presented and a new automatic speakers recognition system is given that performs with 98.9% correct decalcification.
Abstract
n/a
Keywords
Computer Science
Proceedings of the IEEEA tutorial on hidden Markov models and selected applications in speech recognition
22,785 Citations1989L. R. Rabiner
Pattern classification and scene analysis
12,643 Citations1973Richard O. Duda, Peter E. Hart
Discrete-Time Signal Processing
11,537 Citations1989Alan V. Oppenheim, Ronald W. Schafer
The definitive, authoritative text on DSP, written by prominent, DSP pioneers, it provides thorough treatment of the fundamental theorems and properties of discrete- time linear systems, filtering, sampling, and discrete-time Fourier Analysis.
Elsevier eBooksIntroduction to Statistical Pattern Recognition
11,121 Citations1990Keinosuke Fukunaga
A system to determine the sex of fish from their length, as accurately as possible, using a set of labelled data, together with the class label (male or female), as a training set.
Fundamentals of speech recognition
7,687 Citations1993Lawrence R. Rabiner, Biing‐Hwang Juang
This book presents a meta-modelling framework for speech recognition that automates the very labor-intensive and therefore time-heavy and therefore expensive and expensive process of manually modeling speech.
Proceedings of the IEEEOn the use of windows for harmonic analysis with the discrete Fourier transform
7,196 Citations1978F.J. Harris
A comprehensive catalog of data windows along with their significant performance parameters from which the different windows can be compared is included, and an example demonstrates the use and value of windows to resolve closely spaced harmonic signals characterized by large differences in amplitude.
IEEE Transactions on Acoustics Speech and Signal ProcessingDynamic programming algorithm optimization for spoken word recognition
6,544 Citations1978Hiroaki Sakoe, Seibi Chiba
This paper reports on an optimum dynamic progxamming (DP) based time-normalization algorithm for spoken word recognition, in which the warping function slope is restricted so as to improve discrimination between words in different categories.
IEEE ASSP MagazineAn introduction to hidden Markov models
4,789 Citations1986L. R. Rabiner, Biing‐Hwang Juang
The purpose of this tutorial paper is to give an introduction to the theory of Markov models, and to illustrate how they have been applied to problems in speech recognition.
Proceedings of the IEEELinear prediction: A tutorial review
3,996 Citations1975J. Makhoul
This paper gives an exposition of linear prediction in the analysis of discrete signals as a linear combination of its past values and present and past values of a hypothetical input to a system whose output is the given signal.
Journal of the Royal Statistical Society Series A (General)Pattern Classification and Scene Analysis.
3,240 Citations1974Mairi Clarke, Richard O. Duda +1 more
Pattern Recognition Principles
3,205 Citations2009
The present work gives an account of basic principles and available techniques for the analysis and design of pattern processing and recognition systems.
IEEE Transactions on Speech and Audio ProcessingRobust text-independent speaker identification using Gaussian mixture speaker models
2,840 Citations1995D.A. Reynolds, Richard C. Rose
The individual Gaussian components of a GMM are shown to represent some general speaker-dependent spectral shapes that are effective for modeling speaker identity and is shown to outperform the other speaker modeling techniques on an identical 16 speaker telephone speech task.
Pattern RecognitionDigital processing of speech signals
2,647 Citations1980
This paper presents a meta-modelling framework for digital Speech Processing for Man-Machine Communication by Voice that automates the very labor-intensive and therefore time-heavy and expensive process of encoding and decoding speech.
IRE Transactions on Communications SystemsThe Divergence and Bhattacharyya Distance Measures in Signal Selection
1,844 Citations1967T. Kailath
This partly tutorial paper compares the properties of an often used measure, the divergence, with a new measure that is often easier to evaluate, called the Bhattacharyya distance, which gives results that are at least as good and often better than those given by the divergence.
Speech Analysis Synthesis and Perception
1,457 Citations1972James L. Flanagan
Speech CommunicationSpeaker identification and verification using Gaussian mixture speaker models
1,216 Citations1995Douglas A. Reynolds
High performance speaker identification and verification systems based on Gaussian mixture speaker models: robust, statistically based representations of speaker identity, evaluated on four publically available speech databases.
IEEE Transactions on Acoustics Speech and Signal ProcessingCepstral analysis technique for automatic speaker verification
1,170 Citations1981Sadaoki Furui
New techniques for automatic speaker verification using telephone speech based on a set of functions of time obtained from acoustic analysis of a fixed, sentence-long utterance using a new time warping method using a dynamic programming technique.
The Journal of the Acoustical Society of AmericaEffectiveness of linear prediction characteristics of the speech wave for automatic speaker identification and verification
923 Citations1974Bishnu S. Atal
The cepstrum was found to be the most effective, providing an identification accuracy of 70% for speech 50 msec in duration, which increased to more than 98% for a duration of 0.5 sec.
Information Theory and Statistics
741 Citations2005Thomas M. Cover, Joy A. Thomas
Signal ProcessingDistance measures for signal processing and pattern recognition
616 Citations1989Michèle Basseville
Some classical results about error bounds in classification and feature selection for pattern recognition are recalled, which are obtained with the aid of properties of distance measures.
A vector quantization approach to speaker recognition
431 Citations2005Frank K. Soong, A. E. Rosenberg +2 more
A vector quantization (VQ) codebook was used as an efficient means of characterizing the short-time spectral features of a speaker and was used to recognize the identity of an unknown speaker from his/her unlabelled spoken utterances based on a minimum distance (distortion) classification rule.
Proceedings of the IEEEAutomatic recognition of speakers from their voices
374 Citations1976Bishnu S. Atal
The paper indudes a discussion of the speaker-dependent properties of the speech signal, methods for selecting an efficient set of speech measurements, results of experimental studies illustrating the performance of various methods of speaker recognition, and a comparision of theperformance of automatic methods with that of human listeners.
IEEE Signal Processing MagazineText-independent speaker identification
352 Citations1994H. Gish, M. Schmidt
A robust speaker-identification system is presented that was able to deal with various forms of anomalies that are localized in time, such as spurious noise events and crosstalk.
Proceedings of the IEEESpeaker recognition—Identifying people by their voices
345 Citations1985George R. Doddington
A discussion of inherent performance limitations, along with a review of the performance achieved by listening, visual examination of spectrograms, and automatic computer techniques, attempts to provide a perspective with which to evaluate the potential of speaker recognition and productive directions for research into and application of speaker Recognition technology.
IEEE Signal Processing MagazineRobust speaker recognition: a feature-based approach
319 Citations1996Richard J. Mammone, Xiaoyu Zhang +1 more
Linear predictive (LP) analysis, the first step of feature extraction, is discussed, and various robust cepstral features derived from LP coefficients are described, including the afJine transform, which is a feature transformation approach that integrates mismatch to simultaneously combat both channel and noise distortion.
IEEE Transactions on Speech and Audio ProcessingModeling of the glottal flow derivative waveform with application to speaker identification
286 Citations1999Mike Plumpe, Thomas F. Quatieri +1 more
An automatic technique for estimating and modeling the glottal flow derivative source waveform from speech, and applying the model parameters to speaker identification, is presented.
Proceedings of the IEEEAutomatic speaker verification: A review
219 Citations1976A. E. Rosenberg
The techniques, evaluations, and implementations of various proposed speaker recognition systems are reviewed with special emphasis on issues peculiar to speaker verification, especially the distinction between speaker verification and speaker identification.
Testing with the YOHO CD-ROM voice verification corpus
210 Citations2002Joseph P. Campbell
A test plan is presented for the suggested use of the LDC's YOHO CD-ROM for testing voice verification systems, based upon ITT's voice verification test methodology as described by Higgins, et al. (1992).
Digital Signal ProcessingSpeaker verification using randomized phrase prompting
190 Citations1991Alan L. Higgins, L. Bahler +1 more
The system described here is capable of accurately verifying an individual’s claimed identity from a short sample of his or her speech, and a rationale was developed for determining the size of the test required to allow hypotheses regarding the system's true error rates to be tested with stated confidence levels.
IEEE Communications MagazineSpeaker verification: a tutorial
122 Citations1990Jayant M. Naik
The task of speaker verification, a subset of the general problem of speaker recognition, is defined and the feature selection and pattern matching steps of the recognition procedure are examined.
IEEE Transactions on Signal ProcessingOn the application of mixture AR hidden Markov models to text independent speaker recognition
115 Citations1991N.Z. Tisby
The results show that even with a short sequence of only four isolated digits, a speaker can be verified with an average equal-error rate of less than 3 %, and the small improvement over the vector quantization approach indicates the weakness of the Markovian transition probabilities for characterizing speaker-dependent transitional information.
Statistical ScienceDiscriminant Analysis and Clustering: Panel on Discriminant Analysis, Classification, and Clustering
83 Citations1989
Speech CommunicationSpeaker-dependent-feature extraction, recognition and processing techniques
80 Citations1991Sadaoki Furui
Recent advances in and perspectives of research on speaker-dependent-feature extraction from speech waves, automatic speaker identification and verification, speaker adaptation in speech recognition, and voice conversion techniques are discussed.
IEEE Transactions on ComputersOn a New Class of Bounds on Bayes Risk in Multihypothesis Pattern Recognition
70 Citations1974Pierre A. Devijver
A new distance is proposed which permits tighter bounds to be set on the error probability of the Bayesian decision rule and which is shown to be closely related to several certainty or separability measures.
IEEE International Conference on Acoustics Speech and Signal ProcessingVoice identification using nearest-neighbor distance measure
63 Citations1993Alan L. Higgins, Lawrence G. Bahler +1 more
An algorithm for attributing a sample of unconstrained speech to one of several known speakers is described, based on measurement of the similarity of distributions of features extracted from reference speech samples and from the sample to be attributed.
Low-Bit Rate Speech Encoders Based on Line-Spectrum Frequencies (LSFs)
57 Citations1985G. Kang, L. Fransen
This report presents in this report a means for implementing 800- and 4800- b/s voice processors that use a similar speech synthesis filter in which control parameters are line-spectrum frequencies (LSFs).
An approach to text-independent speaker recognition with short utterances
55 Citations2005K. Li, Edwin H. Wrench
A new technique for text-independent speaker recognition is proposed which uses a statistical model of the speaker's vector quantized speech which retains text- independent properties while allowing considerably shorter test utterances than comparable speaker recognition systems.
IEEE Transactions on Acoustics Speech and Signal ProcessingText-independent speaker recognition from a large linguistically unconstrained time-spaced data base
55 Citations1979J. Markel, S. Davis
The application of probability density estimation to text-independent speaker identification
35 Citations2005Richard Schwartz, S. Roucos +1 more
This paper develops the use of probability density function (pdf) estimation for text-independent speaker identification and compares the performance of two parametric and one non-parametric pdf estimation methods to one distance classification method that uses the Mahalanobis distance.
A new method of text-independent speaker recognition
27 Citations2005Alan L. Higgins, R. Wohlford
A new method, based on template matching, that utilizes temporal information to advantage in text-dependent recognition as a special case and is compared with that of similar recently-developed methods.
Cohort selection and word grammar effects for speaker recognition
23 Citations2002John M. Colombi, D.W. Ruck +3 more
This work examines the use of speaker-dependent monophone models to meet the requirements of automatic speaker recognition systems and defines a test hypothesis, a critical error analysis for speaker verification, and a new Bhattacharyya distance for cohort selection.
A TMs32020-based real time, text-independent, automatic speaker verification system
19 Citations2003Joseph B. Attili, M. Savic +1 more
A fast, reliable, yet inexpensive automatic speaker verification system based on the Texas Instruments TMS32020 digital signal processor (DSP) which uses a novel speaker verification algorithm which operates in 75% of real time and requires two to three seconds of unconstrained speech to perform accurate authentication.
The Journal of the Acoustical Society of AmericaText-independent speaker recognition with short utterances
14 Citations1982Kairui Li, Edwin H. Wrench
A new approach to text‐independent speaker recognition, developed to perform with short unknown utterances, models the spectral traits of a speaker with multiple sub‐models rather than using a single statistical distribution as done with previous approaches.
IEEE Transactions on Signal ProcessingInformation-theoretic distortion measures for speech recognition
10 Citations1991Y.-T. Lee
This general framework defines three broad families of distortion measures for speech recognition and provides a consistent way of combining the energy and the spectral information of a phonetic event.
