Using Sparse Representations For Missing Data Imputation In Noise Robust Speech Recognition
Zenodo (CERN European Organization for Nuclear Research)Published 25 August 2008Open access
Cranen, Bert, Gemmeke, Jort
Citations40
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
A novel imputation technique working on entire words that achieves recognition accuracies of 92% at SNR -5 dB using oracle masks on AURORA-2 as compared to 61% using a conventional frame-based approach.
Abstract
Publication in the conference proceedings of EUSIPCO, Lausanne, Switzerland, 2008
Keywords
Computer ScienceEngineering
Journal of the Royal Statistical Society Series B (Statistical Methodology)Regression Shrinkage and Selection Via the Lasso
51,790 Citations1996Robert Tibshirani
A new method for estimation in linear models called the lasso, which minimizes the residual sum of squares subject to the sum of the absolute value of the coefficients being less than a constant, is proposed.
Compressed sensing
17,129 Citations2004David L. Donoho
The Annals of StatisticsLeast angle regression
9,493 Citations2004Bradley Efron, Trevor Hastie +2 more
SIAM Journal on ComputingSparse Approximate Solutions to Linear Systems
2,822 Citations1995B. K. Natarajan
It is shown that the problem is NP-hard, but that the well-known greedy heuristic is good in that it computes a solution with at most at most $\left\lceil 18 \mbox{ Opt} ({\bf \epsilon}/2) \|{\bf A}^+\|^2_2 \ln(\|b\|_2/{\bf
Communications on Pure and Applied MathematicsFor most large underdetermined systems of linear equations the minimal đ<sub>1</sub>ânorm solution is also the sparsest solution
2,541 Citations2006David L. Donoho
The techniques include the use of random proportional embeddings and almostâspherical sections in Banach space theory, and deviation bounds for the eigenvalues of random Wishart matrices.
The aurora experimental framework for the performance evaluation of speech recognition systems under noisy conditions
1,782 Citations2000David J. Pearce, HansâGĂŒnter Hirsch
A database designed to evaluate the performance of speech recognition algorithms in noisy conditions and recognition results are presented for the first standard DSR feature extraction scheme that is based on a cepstral analysis.
Speech CommunicationRobust automatic speech recognition with missing and unreliable acoustic data
593 Citations2001Martin Cooke, Phil Green +2 more
An approach to robust ASR which acknowledges the fact that some spectro-temporal regions will be dominated by noise, and introduces two approaches for dealing with unreliable evidence, including marginalisation and state-based data imputation.
Speech CommunicationSpeech recognition by machines and humans
528 Citations1997Richard P. Lippmann
Comparisons suggest that the human-machine performance gap can be reduced by basic research on improving low-level acoustic-phonetic modeling, on improving robustness with noise and channel variability, and on more accurately modeling spontaneous speech.
Feature Selection in Face Recognition: A Sparse Representation Perspective
165 Citations2007Allen Y. Yang, John Wright +2 more
If sparsity in the recognition problem is properly harnessed, the choice of features is no longer critical and the differences in performance between different features become insignificant as the feature-space dimension is sufficiently large.
Speech CommunicationA Bayesian classifier for spectrographic mask estimation for missing feature speech recognition
152 Citations2004Michael L. Seltzer, Bhiksha Raj +1 more
A new mask estimation technique is presented that uses a Bayesian classifier to determine the reliability of spectrographic elements and resulted in significantly better recognition accuracy than conventional mask estimation approaches.
Speech CommunicationDecoding speech in the presence of other sources
140 Citations2004Jon Barker, Martin Cooke +1 more
This paper proposes a statistical theory of speech recognition in the presence of other acoustic sources by introducing a segregation model in addition to the conventional acoustic and language models and derives an efficient HMM decoder, which searches both across subword state and across alternative segregations of the signal between target and interference.
Towards the generation of French phonetic inflected forms
70 Citations1999Frédérique Sannier, Véronique Aubergé
This paper addresses the problem of identifying reliable regions and proposes two criteria to solve this based on negative energy and SNR and shows that in this task the missing data method performs considerably better than spectral subtraction and the combination of the two techniques outperforms either technique used alone.
Computer Speech & LanguageOn noise masking for automatic missing data speech recognition: A survey and discussion
53 Citations2006Christophe Cerisara, Sébastien Demange +1 more
The objective of this study is to identify the mask estimation methods that have been proposed so far, and to open this domain up to other related research, which could be adapted to overcome this difficult challenge.
Robust speech recognition using cepstral domain missing data techniques and noisy masks
49 Citations2004Hugo Van hamme
A recognizer based on the recently described cepstral-domain MDT approach using missing data masks computed from the noisy signal is described, which exploits a novel decision criterion that integrates harmonicity with signal-to-noise ratio and which makes minimal assumptions on the noise.
Rice Digital Scholarship Archive (Rice University)When is Missing Data Recoverable?
41 Citations2006Yin Zhang
PROSPECT features and their application to missing data techniques for robust speech recognition
31 Citations2004Hugo Van hamme
Alternative to the cepstral representation that lead to more efficient MDT systems are studied, and the proposed solution, PROSPECT features (Projected Spectra), can be interpreted as a novel speech representation, or as an approximation of the inverse covariance matrix of the Gaussian distributions modeling the log-spectra.
Handling Time-Derivative Features in a Missing Data Framework for Robust Automatic Speech Recognition
14 Citations2006Hugo Van hamme
A novel approach to handling dynamic features for automatic speech recognition using a HMM/GMM-architecture and based on missing data techniques for noise robustness is presented.
