Analysis and transcription of general audio data
DSpace@MIT (Massachusetts Institute of Technology)Published 1 January 2000Open access
Michelle S. Spina
Citations3
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
Abstract
Thesis (Ph.D.)--Massachusetts Institute of Technology, Dept. of Electrical Engineering and Computer Science, 2000.
Keywords
Computer Science
Journal of the Royal Statistical Society Series B (Statistical Methodology)Maximum Likelihood from Incomplete Data Via the <i>EM</i> Algorithm
49,657 Citations1977A. P. Dempster, N. M. Laird +1 more
Elements of Information Theory
37,533 Citations2001Thomas M. Cover, Joy A. Thomas
Proceedings of the IEEEA tutorial on hidden Markov models and selected applications in speech recognition
22,785 Citations1989L. R. Rabiner
Pattern classification and scene analysis
12,643 Citations1973Richard O. Duda, Peter E. Hart
Journal of Business and Economic StatisticsAn Introduction to Multivariate Statistical Analysis
9,232 Citations1986Robb J. Muirhead, T. W. Anderson
Fundamentals of speech recognition
7,687 Citations1993Lawrence R. Rabiner, Biing‐Hwang Juang
This book presents a meta-modelling framework for speech recognition that automates the very labor-intensive and therefore time-heavy and therefore expensive and expensive process of manually modeling speech.
Americanae (AECID Library)TIMIT Acoustic-Phonetic Continuous Speech Corpus
2,549 Citations2024John S. Garofolo
Computer Speech & LanguageMaximum likelihood linear regression for speaker adaptation of continuous density hidden Markov models
2,210 Citations1995C.J. Leggetter, Philip C. Woodland
An important feature of the method is that arbitrary adaptation data can be used—no special enrolment sentences are needed and that as more data is used the adaptation performance improves.
SWITCHBOARD: telephone speech corpus for research and development
2,140 Citations1992J. Godfrey, E. Holliman +1 more
The design for the wall street journal-based CSR corpus
1,130 Citations1992Douglas B. Paul, Janet M. Baker
A stochastic parts program and noun phrase parser for unrestricted text
974 Citations1988Kenneth Church
A program that tags each word in an input sentence with the most likely part of speech has been written and performance is encouraging; a 400-word sample is presented and is judged to be 99.5% correct.
IEEE Transactions on Acoustics Speech and Signal ProcessingSpeaker-independent phone recognition using hidden Markov models
936 Citations1989K.-F. Lee, Hsiao-Wuen Hon
The authors introduce the co-occurrence smoothing algorithm, which enables accurate recognition even with very limited training data, and can be used as benchmarks to evaluate future systems.
IEEE MultimediaContent-based classification, search, and retrieval of audio
792 Citations1996Erling Wold, Thom Blum +2 more
The audio analysis, search, and classification engine described here reduces sounds to perceptual and acoustical features, which lets users search or retrieve sounds by any one feature or a combination of them, by specifying previously learned classes based on these features.
Speaker, Environment and Channel Change Detection and Clustering via the Bayesian Information Criterion
680 Citations1998Song M Chen
The segmentation algorithm can successfully detect acoustic changes; the clustering algorithm can produce clusters with high purity, leading to improvements in accuracy through unsupervised adaptation as much as the ideal clustering by the true speaker identities.
International Conference on Acoustics, Speech, and Signal ProcessingSome statistical issues in the comparison of speech recognition algorithms
643 Citations2003Larry Gillick, Simón Cox
The authors present two simple tests for deciding whether the difference in error rates between two algorithms tested on the same data set is statistically significant.
Real-time discrimination of broadcast speech/music
471 Citations2002J. Saunders
A technique which is successful at discriminating speech from music on broadcast FM radio is described, which provides the capability to robustly distinguish the two classes and runs easily in real time.
Proceedings of SPIE, the International Society for Optical Engineering/Proceedings of SPIE<title>Content-based retrieval of music and audio</title>
394 Citations1997Jonathan Foote
A system to retrieve audio documents y acoustic similarity based on statistics derived from a supervised vector quantizer, rather than matching simple pitch or spectral characteristics, which may be applicable to image retrieval as well.
IEEE Signal Processing MagazineText-independent speaker identification
352 Citations1994H. Gish, M. Schmidt
A robust speaker-identification system is presented that was able to deal with various forms of anomalies that are localized in time, such as spurious noise events and crosstalk.
Multimedia SystemsAn overview of audio information retrieval
339 Citations1999Jonathan Foote
The state of the art in audio information retrieval is reviewed, and recent advances in automatic speech recognition, word spotting, speaker and music identification, and audio similarity are presented with a view towards making audio less “opaque”.
Automatic Segmentation, Classification and Clustering of Broadcast News Audio
313 Citations1997M.A. Siegler
This work describes the problems faced in adapting a system built to recognize one utterance at a time to a task that requires recognition of an entire half hour show, and shows that a priori knowledge of acoustic conditions and speakers in the broadcast data is not required for segmentation.
ScholarlyCommons (University of Pennsylvania)A corpus-based approach to language learning
280 Citations1993Eric Brill
A learning algorithm is described that takes a small structurally annotated corpus of text and a larger unannotated corpus as input and automatically learns how to assign accurate structural descriptions to sentences not in the training corpus.
International Conference on Acoustics, Speech, and Signal ProcessingEnvironmental robustness in automatic speech recognition
275 Citations2002Alejandro Acero, Richard M. Stern
Initial efforts to make Sphinx, a continuous-speech speaker-independent recognition system, robust to changes in the environment are reported, and two novel methods based on additive corrections in the cepstral domain are proposed.
Speech CommunicationSubword-based approaches for spoken document retrieval
175 Citations2000Kenney Ng, Victor W. Zue
It is found that with the appropriate subword units, it is possible to achieve performance comparable to that of text-based word units if the underlying phonetic units are recognized correctly.
A probabilistic framework for feature-based speech recognition
129 Citations2002James Glass, J. Chang +1 more
This paper examines a maximum a-posteriori decoding strategy for feature-based recognizers and develops a normalization criterion that is useful for a segment-based Viterbi or A* search.
Segmentation of speech using speaker identification
96 Citations2002Lynn Wilcox, Fankai Chen +2 more
This paper describes techniques for segmentation of conversational speech based on speaker identity using Viterbi decoding on a hidden Markov model network consisting of interconnected speaker sub-networks.
High performance speaker-independent phone recognition using CDHMM
94 Citations1993Lori Lamel, Jean‐Luc Gauvain
It is shown that it is worthwhile to perform phone recognition experiments as opposed to only focusing attention on word recognition results, and high phone accuracies on three corpora: WSJ0, BREF and TIMIT are reported.
Segment generation and clustering in the HTK broadcast news transcription system
92 Citations1998Thomas Hain, SE Johnson +3 more
This paper describes the segmentation, gender detection and segment clustering scheme used in the 1997 HTK broadcast news evaluation system and presents results on both the unpartitioned 1996 development and the 1997 evaluation sets.
Heterogeneous acoustic measurements and multiple classifiers for speech recognition
90 Citations1999Andrew K. Halberstadt, James Glass
Pyrotinib showed activity against NSCLC with HER2 exon 20 mutations in both patient-derived organoids and a PDX model and showed promising efficacy in the clinical trial.
Real-time telephone-based speech recognition in the Jupiter domain
62 Citations1999James Glass, Timothy J. Hazen +1 more
This paper describes the experiences with developing a real-time telephone-based speech recognizer as part of a conversational system in the weather information domain and describes the development of the recognizer vocabulary, pronunciations, language and acoustic models for this system.
Finding Acoustic Regularities in Speech: Applications to Phonetic Recognition
56 Citations1988James Glass
This thesis presents an alternative approach whereby this phonetic-level description is bypassed in favor of directly relating the acoustic realizations to the underlying phonemic forms, and relies critically on the ability to detect important acoustic landmarks in the speech signal.
Automatic transcription of general audio data: preliminary analyses
46 Citations2002Michelle S. Spina, Victor W. Zue
Preliminary analyses and experiments conducted on data collected from a radio news program found that using relatively straightforward acoustic measurements and classification techniques, it was able to achieve better than 80% classification accuracy for seven salient sound classes present in the data, and nearly 94% classified accuracy for a speech/non-speech decision.
Segmentation and modeling in segment-based recognition
32 Citations1997Jane W. Chang, James Glass
The acoustic segmentation algorithm is replaced with “segmentation by recognition,” a probabilistic algorithm that can combine multiple contextual constraints towards hypothesizing only the most likely segments and an efficient search algorithm is described that can efficiently use multiple models to enforce contextual constraints across all segments in a network.
Audio indexing for broadcast news
27 Citations1998S. Dharanipragada, Martin Franz +1 more
The IBM Audio-Indexing System is described which is a combination of a large vocabulary speech recognizer and a text-based information retrieval system that was used to produce the baseline transcripts for the NIST SDR97 evaluation.
Using aggregation to improve the performance of mixture Gaussian acoustic models
25 Citations2002Timothy J. Hazen, A.K. Halberstdt
This paper investigates the use of aggregation as a means of improving the performance and robustness of mixture Gaussian models and produces models that are more accurate and more robust to different test sets than traditional cross-validation using a development set.
Deleted interpolation and density sharing for continuous hidden Markov models
21 Citations2002Xuedong Huang, Mei-Yuh Hwang +2 more
This paper proposes to smooth the probability density values instead of the parameters of continuous HMMs, and points out that deleted interpolation can be regarded as a parameter sharing technique that reduced the word error rate over other simple parameter smoothing techniques.
DSpace@MIT (Massachusetts Institute of Technology)Near-miss modeling : a segment-based approach to speech recognition
18 Citations1998Jane W. Chang
This thesis describes an approach called near-miss modeling that addresses the major difficulties in segment-based recognition and runs a recognizer and produces a graph that contains only the segments on paths that score within a threshold of the best scoring path.
Cambridge University Engineering Department Publications DatabaseIBM's LVCSR system for transcription of broadcast news used in the 1997 hub4 english evaluation
16 Citations1998S Chen, Mark Gales +5 more
IBM’s large vocabulary continuous speech recognition (LVCSR) system used in the 1997 Hub4 English evaluation includes a number of new features: optimal feature space for acoustic modeling (in training and/or testing), filler-word modeling, Bayesian Information Criterion (BIC) based segmentation and segment clustering, and 4-gram language models.
A model distance measure for talker clustering and identification
15 Citations2002Jonathan Foote, H.F. Silverman
This paper describes methods of talker clustering and identification based on a "distance" metric between discrete HMM output probabilities derived on a tree-based MMI partition of the feature space, rather than the usual vector quantization.
Automatic transcription of general audio data: effect of environment segmentation on phonetic recognition 1
6 Citations1997Michelle S. Spina, Victor W. Zue
This work has studied the effects of different speaking environments on a phonetic recognition task using data collected from a radio news program, and found that if a singlerecognizer is to be used, it is more effective to use a smaller amount of homogeneous, clean data for training.
Toward Content-Based Audio Indexing and Retrieval and a New Speaker Discrimination Technique
5 Citations1995Lonce Wyse, Stephen W. Smoliar
