Sensor fusion weighting measures in Audio-Visual Speech Recognition
Published 1 January 2004
Trent Lewis, David Powers
Citations23
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
This paper analyses current weighting measures and compares them to several new measures proposed by the authors, finding that when calculating the dispersion of the output there is a shift from analyzing the variance to analysing the skewness of the distribution.
Abstract
All in-text\treferences\tunderlined\tin\tblue\tare\tlinked\tto\tpublications\ton\tResearchGate, letting you\taccess\tand\tread\tthem\timmediately.
Keywords
Computer Science
Fundamentals of speech recognition
7,687 Citations1993Lawrence R. Rabiner, Biing‐Hwang Juang
This book presents a meta-modelling framework for speech recognition that automates the very labor-intensive and therefore time-heavy and therefore expensive and expensive process of manually modeling speech.
NatureHearing lips and seeing voices
6,130 Citations1976Harry McGurk, John MacDonald
The study reported here demonstrates a previously unrecognised influence of vision upon speech perception, on being shown a film of a young woman's talking head in which repeated utterances of the syllable [ba] had been dubbed on to lip movements for [ga].
IEEE Transactions on Acoustics Speech and Signal ProcessingComparison of parametric representations for monosyllabic word recognition in continuously spoken sentences
5,322 Citations1980S. Davis, P. Mermelstein
Several parametric representations of the acoustic signal were compared with regard to word recognition performance in a syllable-oriented continuous speech recognition system and the superior performance of the mel-frequency cepstrum coefficients may be attributed to the fact that they better represent the perceptually relevant aspects of the short-term speech spectrum.
An Introduction to Language
1,093 Citations1974Victoria A. Fromkin, Robert D. Rodman +1 more
IEEE Transactions on MultimediaAudio-visual speech modeling for continuous speech recognition
597 Citations2000Stéphane Dupont, Juergen Luettin
A speech recognition system that uses both acoustic and visual speech information to improve recognition performance in noisy environments and is demonstrated on a large multispeaker database of continuously spoken digits.
Proceedings of the IEEEAudio-visual integration in multimodal communication
251 Citations1998Tsuhan Chen, R.R. Rao
This work reviews recent research that examines audio-visual integration in multimodal communication, including bimodality in human speech, human and automated lip reading, facial animation, lip synchronization, joint audio-video coding, andbimodal speaker verification.
IEEE Transactions on Evolutionary ComputationDesigning classifier fusion systems by genetic algorithms
240 Citations2000Lakhmi C. Jain, Ludmila I. Kuncheva
Two simple ways to use a genetic algorithm (GA) to design a multiple-classifier system are suggested that can be made less prone to overtraining by including penalty terms in the fitness function accounting for the number of features used.
Journal of Speech and Hearing ResearchEffects of Training on the Visual Recognition of Consonants
220 Citations1977Brian E. Walden, Robert A. Prosek +3 more
Visual recognition of consonants was studied in 31 hearing-impaired adults before and after 14 hours of concentrated, individualized, spechreading training and revealed that most changes in consonant recognition occurred during the first few hours of training.
NATO ASI series. Series F : Computer and system sciencesOn the Integration of Auditory and Visual Parameters in an HMM-based ASR
160 Citations1996A. Adjoudani, Christian Benoı̂t
A model which can improve the performances of an audio-visual speech recognizer in an isolated word and speaker dependent situation is proposed by using a hybrid system based on two HMMs trained respectively with acoustic and optic data.
NATO ASI series. Series F : Computer and system sciencesVisionary Speech: Looking Ahead to Practical Speechreading Systems
106 Citations1996Marcus E. Hennecke, David G. Stork +1 more
An overview of speechreading systems is presented, paying particular attention to the various approaches to key design decisions and the benefits and drawbacks of each.
Large-vocabulary audio-visual speech recognition: a summary of the Johns Hopkins Summer 2000 Workshop
92 Citations2001C. Neti, Gerasimos Potamianos +4 more
In a summary of the Johns Hopkins Summer 2000 Workshop on audio-visual automatic speech recognition (ASR), introducing the visual modality reduced ASR word error rate by 7% relative in clean speech, and by 27% relative at an 8.5 dB SNR audio condition.
Weighting schemes for audio-visual fusion in speech recognition
80 Citations2001Hervé Glotin, D. Vergyr +3 more
An improvement in the state-of-the-art large vocabulary continuous speech recognition (LVCSR) performance is demonstrated by the use of visual information, in addition to the traditional audio one, by taking a decision fusion approach for the audio-visual information.
Adaptive bimodal sensor fusion for automatic speechreading
76 Citations2002Uwe Meier, Wolfgang Hürst +1 more
Different methods of combining the visual and acoustic data to improve the recognition performance of automated speech recognizers by using additional visual information are presented, achieving error reduction of up to 50%.
Journal of Intelligent & Robotic SystemsSensor data fusion
71 Citations1988L. F. Pau
This paper reviews some knowledge representation approaches devoted to the sensor fusion problem, as encountered whenever images, signals, text must be combined to provide the input to a controller or to an inference procedure.
International Journal of Pattern Recognition and Artificial IntelligenceTOWARDS UNRESTRICTED LIP READING
62 Citations2000Uwe Meier, Rainer Stiefelhagen +2 more
The experimental results indicate that the system can achieve up to 55% error reduction using visual information in addition to the acoustic signal, and the feasibility of the proposed methods is demonstrated by the development of a modular system for flexible human–computer interaction via both visual and acoustic speech.
Lecture notes in computer scienceA Framework for Classifier Fusion: Is It Still Needed?
50 Citations2000Josef Kittler
This work considers the problem and issues of classifier fusion and adopts the Bayesian viewpoint and shows how this leads to classifier output moderation to compensate for sampling problems, and elaborate how the final stage of fusion should combine the complementary measurement information that might be available to different experts.
Audio-visual interaction in multimedia communication
35 Citations2002Tsuhan Chen, R.R. Rao
This paper presents the results in exploiting the audio-visual interaction that is very significant in multimedia communication, including lip synchronization, joint audio-video coding, and person verification.
Machine LearningRobust Sensor Fusion: Analysis and Application to Audio Visual Speech Recognition
30 Citations1998Javier R. Movellan, Paul Mineiro
This paper proposes a principled solution to the issue of catastrophic fusion in multimodal recognition systems that integrate the output from several modules while working in non-stationary environments based upon Bayesian ideas of competitive models and inference robustification.
The learning behavior of single neuron classifiers on linearly separable or nonseparable input
29 Citations2003Mitra Basu, Tin Kam Ho
This work explores the behavior of several classical descent procedures for determining linear separability and training linear classifiers in the presence of linearly non separated input and finds that the adaptive procedures have serious implementation problems which make them less preferable than linear programming.
Lecture notes in computer scienceContinuous audio-visual speech recognition
27 Citations1998Juergen Luettin, Stéphane Dupont
This work tackles the problem of joint temporal modelling of the acoustic and visual speech signals by applying Multi-Stream hidden Markov models and allows the use of different temporal topologies and levels of stream integration and hence enables to model temporal dependencies more accurately.
Cross-modal prediction in audio-visual communication
23 Citations2002R.R. Rao, Tsuhan Chen
A novel means for predicting the shape of a person's mouth from the corresponding speech signal is presented and applications of this prediction to video coding are explored.
International Journal of Artificial Intelligence ToolsDISCRIMINATIVE LEARNING OF VISUAL DATA FOR AUDIOVISUAL SPEECH RECOGNITION
21 Citations1999Alexandrina Rogozan
Methods for reinforcing the visible speech recognition in the framework of separate identification are outlined and it is shown that using these methods improves performances of the DI+SI based system under varying noise-level conditions.
A hybrid ANN/HMM audio-visual speech recognition system.
21 Citations2001Martin Heckmann, Frédéric Berthommier +1 more
A system for audio-visual speech recognition based on a hybrid Artificial Neural Network/Hidden Markov Model (ANN/HMM) approach and an implementation of this method in an automatic fusion depending on the noise level in the audio channel is developed.
Proceedings of SPIE, the International Society for Optical Engineering/Proceedings of SPIE<title>Preprocessing video images for neural learning of lipreading</title>
20 Citations1994Kislaya Prasad, David G. Stork +1 more
New procedures for preprocessing video images for automatic lipreading applications are presented and it is shown that these procedures can be used for improving the quality of lipreading application results.
Optimal weighting of posteriors for audio-visual speech recognition
17 Citations2002Martin Heckmann, F. Berthommier +1 more
This work investigates the fusion of audio and video a posteriori phonetic probabilities in a hybrid ANN/HMM audio-visual speech recognition system and compares these two new concepts in audio- visual recognition to a rather standard approach known from the literature.
NATO ASI series. Series F : Computer and system sciencesTowards a Robust Speechreading Dialog System
8 Citations1996Christoph Bregler, Stephen M. Omohundro +2 more
A hybrid speechreading system that is based on a Manifold Learning technique, on Neural Networks, and on Hidden Markov Models significantly improves the performance of acoustic speech recognizers in degraded environments.
