Speech emotion recognition using hidden Markov models
Speech CommunicationPublished 1 August 2003Open access
Tin Lay Nwe, Say Wei Foo, Liyanage C. De Silva
Citations907
SJR quartileQ4
SJR score0.11
SNIP0.06
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
This paper proposes a text independent method of emotion classification of speech that makes use of short time log frequency power coefficients (LFPC) to represent the speech signals and a discrete hidden Markov model (HMM) as the classifier.
Abstract
10.1016/S0167-6393(03)00099-2
Keywords
PsychologyComputer Science
John Murray eBooksThe expression of the emotions in man and animals.
11,390 Citations1872Charles Darwin
Fundamentals of speech recognition
7,687 Citations1993Lawrence R. Rabiner, Biing‐Hwang Juang
This book presents a meta-modelling framework for speech recognition that automates the very labor-intensive and therefore time-heavy and therefore expensive and expensive process of manually modeling speech.
IEEE Transactions on Acoustics Speech and Signal ProcessingComparison of parametric representations for monosyllabic word recognition in continuously spoken sentences
5,322 Citations1980S. Davis, P. Mermelstein
Several parametric representations of the acoustic signal were compared with regard to word recognition performance in a syllable-oriented continuous speech recognition system and the superior performance of the mel-frequency cepstrum coefficients may be attributed to the fact that they better represent the perceptually relevant aspects of the short-term speech spectrum.
Pattern RecognitionDigital processing of speech signals
2,647 Citations1980
This paper presents a meta-modelling framework for digital Speech Processing for Man-Machine Communication by Voice that automates the very labor-intensive and therefore time-heavy and expensive process of encoding and decoding speech.
IEEE Signal Processing MagazineEmotion recognition in human-computer interaction
2,507 Citations2001Roddy Cowie, Ellen Douglas‐Cowie +5 more
Basic issues in signal processing and analysis techniques for consolidating psychological and linguistic analyses of emotion are examined, motivated by the PKYSTA project, which aims to develop a hybrid system capable of using information from faces and voices to recognize people's emotions.
Discrete-Time Processing of Speech Signals
2,462 Citations1999J.R. Deller, John H. L. Hansen +1 more
The preface to the IEEE Edition explains the background to speech production, coding, and quality assessment and introduces the Hidden Markov Model, the Artificial Neural Network, and Speech Enhancement.
Psychological BulletinVocal affect expression: A review and a model for future research.
1,636 Citations1986Klaus R. Scherer
A "component patterning" model of vocal affect expression is proposed that attempts to rink the outcomes of antecedent event evaluation to biologically based response patterns and may help to stimulate hypothesis-guided research as well as provide a framework for the development of appropriate research paradigms.
Cambridge University Press eBooksThe Expression of the Emotions in Man and Animals
1,353 Citations2013Charles Darwin
The Journal of the Acoustical Society of AmericaToward the simulation of emotion in synthetic speech: A review of the literature on human vocal emotion
1,093 Citations1993Iain R. Murray, John L. Arnott
The voice parameters affected by emotion are found to be of three main types: voice quality, utterance timing, and utterance pitch contour.
IEEE Transactions on Acoustics Speech and Signal ProcessingSpeaker-independent phone recognition using hidden Markov models
936 Citations1989K.-F. Lee, Hsiao-Wuen Hon
The authors introduce the co-occurrence smoothing algorithm, which enables accurate recognition even with very limited training data, and can be used as benchmarks to evaluate future systems.
The Journal of the Acoustical Society of AmericaEffectiveness of linear prediction characteristics of the speech wave for automatic speaker identification and verification
923 Citations1974Bishnu S. Atal
The cepstrum was found to be the most effective, providing an identification accuracy of 70% for speech 50 msec in duration, which increased to more than 98% for a duration of 0.5 sec.
Journal of Cross-Cultural PsychologyEmotion Inferences from Vocal Expression Correlate Across Languages and Cultures
743 Citations2001Klaus R. Scherer, Rainer Banse +1 more
The psychology and biology of emotion
725 Citations1994Robert Plutchik
This book focuses on the study of human emotion and the role that emotion plays in the development of a person's personality.
ITL Review of Applied LinguisticsProsodic Systems and Intonation in English
713 Citations1969René Collier
Language<b>Prosodic systems and intonation in English</b> . By David Crystal. Cambridge: University Press, 1969. Pp. viii, 381. $16.00.
708 Citations1976Philip Lieberman
This chapter discusses voice-quality and sound attributes in prosodic study, the intonation system of English, and the semantics ofintonation.
Recognizing emotion in speech
498 Citations2002Frank Dellaert, Thomas Polzin +1 more
A new method of extracting prosodic features from speech, based on a smoothing spline approximation of the pitch contour, is presented, which obtains classification performance that is close to human performance on the task.
Speech CommunicationDigital speech processing, synthesis, and recognition
416 Citations1989
This paper presents principal characteristics of speech speech production models speech analysis and analysis-synthesis systems linear predictive coding (LPC) analysis speech coding speech synthesis speech recognition future directions of speech processing.
IEEE Transactions on Acoustics Speech and Signal ProcessingA new vector quantization clustering algorithm
405 Citations1989W. Equitz
The pairwise nearest neighbor (PNN) algorithm is presented as an alternative to the Linde-Buzo-Gray (1980, LBG) (generalized Lloyd, 1982) algorithm for vector quantization clustering.
American PsychologistIf it's not left, it's right: Electroencephalograph asymmetry and the development of emotion.
373 Citations1991Nathan A. Fox
Medical Entomology and ZoologyThe Science of Emotion: Research and Tradition in the Psychology of Emotion
302 Citations1995Randolph R. Cornelius
Speech MonographsAn experimental study of the pitch characteristics of the voice during the expression of emotion∗
231 Citations1939Grant Fairbanks, Wilbert Pronovost
Medical Entomology and ZoologySpeech Recognition: Theory and C++ Implementation
200 Citations1999C. Becchetti, L. Prina Ricotti
HMM Training.
APPROACHING AUTOMATIC RECOGNITION OF EMOTION FROM VOICE: A ROUGH BENCHMARK
167 Citations2000Sinéad McGilloway, Roddy Cowie +4 more
Characteristics and Recognizability of Vocal Expressions of Emotion
165 Citations1984Renée van Bezooijen
The Journal of the Acoustical Society of AmericaNonlinear analysis and classification of speech under stressed conditions
116 Citations1994Douglas A. Cairns, John H. L. Hansen
It is hypothesize that speech consists of a linear and nonlinear component, and that the non linear component changes markedly between normal and stressed speech.
Archives of General PsychiatryRecognition of Emotion From Vocal Cues
104 Citations1986William F. Johnson
Voice-synthesized samples seemed to capture some cues promoting emotion recognition, but correct identification did not approach that of other segments, and recognition of emotion decreased, but not as dramatically as expected in each of the three alterations of the original samples.
Language and SpeechA New Method of Investigating the Perception of Prosodic Features
101 Citations1978Iván Fónagy
The laryngograph may help to reveal prosodic features conveying emotive attitudes, or characterizing different verbal genres, and proves to be a useful tool in teaching intonation patterns of foreign languages.
Journal of Speech and Hearing ResearchRelations Between Prosodic Variables and Emotions in Normal American English Utterances
93 Citations1968George L. Huttar
Emotion recognition in speech using neural networks
71 Citations2003Joy Nicholson, Kazutoshi Takahashi +1 more
Pattern recognition of emotion with neural network
30 Citations2002Takeshi Yamada, H. Hashimoto +1 more
An emotion model for communication which also transfers personality and character information is proposed, and the emotion model as network agent between two human communication partners is discussed.
