The 2004 BBN/LIMSI 10xRT English Broadcast News Transcription System
HAL (Le Centre pour la Communication Scientifique Directe)Published 1 January 2004
Long Nguyen, Sherif Abdou, Mohamed Afify, John Makhoul, Spyros Matsoukas, Richard Schwartz
Citations24
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
This paper describes the 2004 BBN/LIMSI 10xRT English Broadcast News (BN) transcription system which uses a tightly integrated combination of components from the BBN and LIMSI speech recognition systems, obtaining a word hypothesis that is better than is produced by either system alone, while remaining within the allotted time limit.
Abstract
International audience
Keywords
Computer Science
The Journal of the Acoustical Society of AmericaPerceptual linear predictive (PLP) analysis of speech
2,539 Citations1990Hynek Heřmanský
A new technique for the analysis of speech, the perceptual linear predictive (PLP) technique, which uses three concepts from the psychophysics of hearing to derive an estimate of the auditory spectrum, and yields a low-dimensional representation of speech.
Computer Speech & LanguageMaximum likelihood linear regression for speaker adaptation of continuous density hidden Markov models
2,210 Citations1995C.J. Leggetter, Philip C. Woodland
An important feature of the method is that arbitrary adaptation data can be used—no special enrolment sentences are needed and that as more data is used the adaptation performance improves.
Computer Speech & LanguageMaximum likelihood linear transformations for HMM-based speech recognition
1,490 Citations1998Mark Gales
The paper compares the two possible forms of model-based transforms: unconstrained, where any combination of mean and variance transform may be used, and constrained, which requires the variance transform to have the same form as the mean transform.
A post-processing system to yield reduced word error rates: Recognizer Output Voting Error Reduction (ROVER)
1,093 Citations2002J. Fiscus
A post-recognition process which models the output generated by multiple ASR systems as independent knowledge sources that can be combined and used to generate an output with reduced error rate.
Speech CommunicationThe LIMSI Broadcast News transcription system
374 Citations2002Jean‐Luc Gauvain, Lori Lamel +1 more
Development work in moving from laboratory read speech data to real-world or `found' speech data in preparation for the DARPA evaluations on this task from 1996 to 1999 is described.
Speech CommunicationHeteroscedastic discriminant analysis and reduced rank HMMs for improved speech recognition
339 Citations1998Nagendra Kumar, Andreas G. Andreou
Theoretical results to the problem of speech recognition are applied and word-error reduction in systems that employed both diagonal and full covariance heteroscedastic Gaussian models tested on the TI-DIGITS database is observed.
Computer Speech & LanguageLarge scale discriminative training of hidden Markov models for speech recognition
309 Citations2002Philip C. Woodland, Daniel Povey
It is shown that HMMs trained with MMIE benefit as much as MLE-trained HMMs from applying model adaptation using maximum likelihood linear regression (MLLR), which has allowed the straightforward integration of MMIe- trained HMMs into complex multi-pass systems for transcription of conversational telephone speech.
Computer Speech & LanguageLightly supervised and unsupervised acoustic model training
261 Citations2002Lori Lamel, Jean‐Luc Gauvain +1 more
Experiments providing supervision only via the language model training materials show that including texts which are contemporaneous with the audio data is not crucial for success of the approach, and that the acoustic models can be initialized with as little as 10 min of manually annotated data.
Finding consensus among words: lattice-based word error minimization
217 Citations1999Lidia Mangu, Eric Brill +1 more
A new algorithm for finding the hypothesis in a recognition lattice that is expected to minimize the word error rate (WER) is described, which overcomes the mismatch between the word-based performance metric and the standard MAP scoring paradigm that is sentence-based.
2003 IEEE International Conference on Acoustics, Speech, and Signal Processing, 2003. Proceedings. (ICASSP '03).Online speaker clustering
81 Citations2003D. Liu, Francis Kubala
A set of new algorithms that perform speaker clustering in an online fashion that enables low-latency incremental speaker adaptation in online speech-to-text systems and gives a speaker tracking and indexing system the ability to label speakers with cluster ID on the fly.
Light supervision in acoustic model training
49 Citations2004Long Nguyen, Bing Xiang
A new light supervision method to derive additional acoustic training data automatically for broadcast news transcription systems using a biased language model in the lightly supervised decoding of a subset of the TDT corpus.
Speech recognition in multiple languages and domains: the 2003 BBN/LIMSI EARS system
38 Citations2004Richard Schwartz, Thomas Colthurst +19 more
It is demonstrated that a joint BBN/LIMSI system with a time constraint achieved better results than either system alone.
Neural network language models for conversational speech recognition
33 Citations2004Holger Schwenk, Jean‐Luc Gauvain
The generalization behavior of the neural network LM for in-domain training corpora varying from 7M to over 21M words is analyzed and significant word error reductions were observed compared to a carefully tuned 4-gram backoff language model in a state of the art conversational speech recognizer for the NIST rich transcription evaluations.
Improved speaker adaptation using speaker dependent feature projections
23 Citations2004Spyros Matsoukas, Richard Schwartz
Results on the broadcast news corpus show that the proposed HLDA adaptation technique is very effective, even when combined with traditional CMLLR and MLLR adaptation, providing up to 8% relative improvement in recognition accuracy.
Speech CommunicationProgress in transcription of Broadcast News using Byblos
22 Citations2002Long Nguyen, Spyros Matsoukas +4 more
Besides improving recognition accuracy, the BBN Byblos speech recognition system succeeded in developing several algorithms to achieve close-to-real-time recognition speed without a significant sacrifice in recognition accuracy.
Single-tree method for grammar-directed search
22 Citations1999Long Nguyen, Richard Schwartz
A novel fast-match algorithm that has high-accuracy recognition and run-time proportional to only the cube root of the vocabulary size and is able to use a word bigram language model without making copies of the tree during the search.
