Channel Separability in the Audio-Visual Integration of Speech: A Bayesian Approach
Generate an AI Snapshot to get a quick, structured summary of this paper.
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
Results are presented suggesting that the opto-acoustic signals are indeed conditionally independent and that therefore the factorization of optic and acoustic influences observed in humans is optimal.
Abstract
Experiments show that human perceptual responses to audiovisual speech signals factorize into two independent components, one controlled by the optic signal and one controlled by the acoustic signal (Massaro, 1987). From a Bayesian point of view, this result indicates that, at some level, the perceptual system treats acoustic and optic speech signals as if they were conditionally independent processes. This raises the question of whether conditional independence is an optimal assumption or whether the perceptual system uses it for reasons other than minimization of error rates. In this paper we present results suggesting that the opto-acoustic signals are indeed conditionally independent and that therefore the factorization of optic and acoustic influences observed in humans is optimal. Finally, based on a previous analysis by Movellan and McClelland (1995) we show that the implicit assumption of conditional independence can be implemented in nervous systems by using physically separable audio and visual channels that talk to each other via top-down feedback connections.
