An evaluation of the robustness of existing supervised machine learning approaches to the classification of emotions in speech
Speech CommunicationPublished 2 February 2007Open access
Mohammad Shami, Werner Verhelst
Citations131
SJR quartileQ1
SJR score0.49
SNIP1.25
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
The robustness of approaches to the automatic classification of emotions in speech is addressed and it is suggested that existing approaches are efficient enough to handle larger amounts of training data without any reduction in classification accuracy.
Abstract
An evaluation of the robustness of existing supervised machine learning approaches to the classification of emotions in speech
Keywords
PsychologyComputer Science
The Adapted mind : evolutionary psychology and the generation of culture
5,267 Citations1992Jerome H. Barkow, Leda Cosmides +1 more
PsycEXTRA DatasetAffective computing
5,051 Citations1997Rosalind W. Picard
Key issues in affective computing, " computing that relates to, arises from, or influences emotions", are presented and new applications are presented for computer-assisted learning, perceptual information retrieval, arts and entertainment, and human health and interaction.
ACCURATE SHORT-TERM ANALYSIS OF THE FUNDAMENTAL FREQUENCY AND THE HARMONICS-TO-NOISE RATIO OF A SAMPLED SOUND
1,043 Citations1993Paul Boersma
A straightforward and robust algorithm for periodicity detection, working in the lag (autocorrelation) domain, that is several orders of magnitude more accurate than the methods commonly used for speech analysis and capable of measuring harmonics-to-noise ratios in thelag domain with an accuracy and reliability much greater than that of any of the usual frequency-domain methods.
Speech CommunicationSpeech emotion recognition using hidden Markov models
907 Citations2003Tin Lay Nwe, Say Wei Foo +1 more
This paper proposes a text independent method of emotion classification of speech that makes use of short time log frequency power coefficients (LFPC) to represent the speech signals and a discrete hidden Markov model (HMM) as the classifier.
Hidden Markov model-based speech emotion recognition
581 Citations2003Björn W. Schuller, Gerhard Rigoll +1 more
The paper addresses the design of working recognition engines and results achieved with respect to the alluded alternatives and describes a speech corpus consisting of acted and spontaneous emotion samples in German and English language.
Human Maternal Vocalizations to Infants as Biologically Relevant Signals: An Evolutionary Perspective
490 Citations1992Anne Fernald
International Journal of Human-Computer StudiesThe production and recognition of emotions in speech: features and algorithms
438 Citations2003Oudeyer Pierre-Yves
A technique which allows to continuously control both the age of a synthetic voice and the quantity of emotions that are expressed and the first large-scale data mining experiment about the automatic recognition of basic emotions in informal everyday short utterances is presented.
Speech CommunicationHow to find trouble in communication
281 Citations2003Anton Batliner, Kerstin Fischer +3 more
The module Monitoring of User State [especially of] Emotion (MOUSE) is proposed in which a prosodic classifier is combined with other knowledge sources, such as conversationally peculiar linguistic behavior, for example, the use of repetitions.
Autonomous RobotsRecognition of Affective Communicative Intent in Robot-Directed Speech
245 Citations2002Cynthia Breazeal, Lijin Aryananda
This paper presents an approach for recognizing four distinct prosodic patterns that communicate praise, prohibition, attention, and comfort to preverbal infants and integrates this perceptual ability into the authors' robot's “emotion” system, thereby allowing a human to directly manipulate the robot's affective state.
Spontaneous speech: how people really talk and why engineers should care
168 Citations2005Elizabeth Shriberg
An overview of four fundamental properties of spontaneous speech that present challenges for spoken language applications because they violate assumptions often applied in automatic processing technology are described.
Speaker Independent Speech Emotion Recognition by Ensemble Classification
147 Citations2005Björn W. Schuller, S.A. Reiter +4 more
This work strives to recognize emotion independent of the person concentrating on the speech channel, by addressing single feature relevance of acoustic features is a critical point by filter-based gain ratio calculation and applying an SVM-SFFS wrapper based search.
Automatic Speech Classification To Five Emotional States Based On Gender Information
106 Citations2004Dimitrios Ververidis, Constantine Kotropoulos
The Sequential Forward Selection method (SFS) has been used in order to discover the 5-10 features which are able to classify the samples in the best way for each gender, and a random classification would result in a correct classification rate of 20%.
Speech CommunicationBabyEars: A recognition system for affective vocalizations
104 Citations2003Malcolm Slaney, Gerald W. McRoberts
Mothers' speech was significantly easier to classify than fathers' speech, suggesting either clearer distinctions among these messages in mothers' speech to infants, or a difference between fathers and mothers in the acoustic information used to convey these messages.
Child DevelopmentA Combination of Vocal f0 Dynamic and Summary Features Discriminates between Three Pragmatic Categories of Infant-Directed Speech
103 Citations1996Gary S. Katz, Jeffrey F. Cohn +1 more
To assess the relative contribution of dynamic and summary features of vocal fundamental frequency (f0) to the statistical discrimination of pragmatic categories in infant-directed speech, 49 mothers were instructed to use their voice to get their 4-month-old baby's attention, show approval, and provide comfort.
Speech CommunicationEmotions, speech and the ASR framework
100 Citations2003Louis ten Bosch
The conclusion is that automatic emotional tagging of the speech signal is difficult to perform with high accuracy, but prosodic information is nevertheless potentially useful to improve the dialogue handling in ASR tasks on a limited domain.
Segment-Based Approach to the Recognition of Emotions in Speech
65 Citations2005Mohammad Shami, Mohamed S. Kamel
A new framework for the context and speaker independent recognition of emotions from voice, based on a richer and more natural representation of the speech signal, is proposed, yielding an overall classification accuracy of 87% for 5 emotions, outperforming previous results on a similar database.
Classical and novel discriminant features for affect recognition from speech
63 Citations2005Raul Castro Fernandez, Rosalind W. Picard
The performance and relevance of a set of acoustic features for the task of automatic recognition of affect from speech using machine learning techniques are investigated and it is shown that the more exploratory and novel subset of these features outrank the more classical features in the recognition task.
Research Commons (University of Waikato)Applying propositional learning algorithms to multi-instance data
48 Citations2003Eibe Frank, Xin Xu
A simple wrapper for applying standard propositional learners to multi-instance problems is introduced and it is shown that these two modifications are essential for producing good results on the Musk benchmark.
Tales of tuning - prototyping for automatic classification of emotional user states
39 Citations2005Anton Batliner, Stefan Steidl +3 more
This work presents a database with emotional children’s speech in a human-robot scenario and discusses possible strategies for tuning, e.g., using only prototypes, or taking into account requirements and feasibility in possible applications.
On integrating insights from human speech perception into automatic speech recognition
27 Citations2005Sorin Dusan, L. R. Rabiner
A comparison between HSP and ASR is presented emphasizing some insights from HSP that could still be applied in ASR and some ideas for extracting useful non-linguistic information from the speech signal, the so called ‘rich transcription’, which could help in selecting specialized acoustic-lingUistic models that offer higher accuracy than the general models.
Using word-level pitch features to better predict student emotions during spoken tutoring dialogues
21 Citations2005Mihai Rotaru, Diane Litman
A first comparison of the two levels for the task of predicting student emotions in twocorpora of spoken tutoring dialogues, finding that the advantage of word-level features lies in a betterprediction of longer turns.
Zenodo (CERN European Organization for Nuclear Research)Passive Versus Active: Vocal Classification System
11 Citations2005Zakia Hammal, Barış Bozkurt +4 more
This work proposes to define two classes of expression: Active gathering Happiness, Surprise and Anger versus Passive gathering Neutral and Sadness, and tests several classification methods, namely a Bayesian classifier, a Linear Discriminant Analysis (LDA), the K Nearest Neighbours (KNN) and a Support Vector Machine with gaussian radial basis function kernel (SVM).
