Relevance of time–frequency features for phonetic and speaker-channel classification
Generate an AI Snapshot to get a quick, structured summary of this paper.
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
A large database of hand-labeled fluent speech is used to compute the mutual information between a phonetic classification variable and one spectral feature variable in the time–frequency plane, and the joint mutual information (JMI) between the phonetic Classification variable and two feature variables in thetime-frequency plane.
Abstract
The mutual information concept is used to study the distribution of speech information in frequency and in time. The main focus is on the information that is relevant for phonetic classification. A large database of hand-labeled fluent speech is used to (a) compute the mutual information (MI) between a phonetic classification variable and one spectral feature variable in the time–frequency plane, and (b) compute the joint mutual information (JMI) between the phonetic classification variable and two feature variables in the time–frequency plane. The MI and the JMI of the feature variables are used as relevance measures to select inputs for phonetic classifiers. Multi-layer perceptron (MLP) classifiers with one or two inputs are trained to recognize phonemes to examine the effectiveness of the input selection method based on the MI and the JMI. To analyze the non-linguistic sources of variability, we use speaker-channel labels to represent different speakers and different telephone channels and estimate the MI between the speaker-channel variable and one or two feature variables.
