Abstract
We summarize recent progress in automatic specch-to-text transcription at SRI, ICSI, and the University of Washington. The work encompasses all components of speech modeling found in a state-of-the-art recognition system, from acoustic features, to acoustic modeling and adaptation, to language modeling. In the front end, we experimented with nonstandard features, including various measures of voicing, discriminative phone posterior features estimated by multilayer perceptrons, and a novel phone-level macro-averaging for cepstral normalization. Acoustic modeling was improved with combinations of front ends operating at multiple frame rates, as well as by modifications to the standard methods for discriminative Gaussian estimation. We show that acoustic adaptation can be improved by predicting the optimal regression class complexity for a given speaker. Language modeling innovations include the use of a syntax-motivated almost-parsing language model, as well as principled vocabulary-selection techniques. Finally, we address portability issues, such as the use of imperfect training transcripts, and language-specific adjustments required for recognition of Arabic and Mandarin.
| Original language | English (US) |
|---|---|
| Pages (from-to) | 1729-1742 |
| Number of pages | 14 |
| Journal | IEEE Transactions on Audio, Speech and Language Processing |
| Volume | 14 |
| Issue number | 5 |
| DOIs | |
| State | Published - Sep 2006 |
| Externally published | Yes |
Funding
Manuscript received October 16, 2005; revised May 30, 2006. This work was supported by the Defense Advanced Research Projects Agency (DARPA) under Contract MDA972-02-C-0038 and Grant MDA972-02-1-0024 (approved for public release, distribution unlimited). The associate editor coordinating the review of this manuscript and approving it for publication was Dr. Alex Acero.
| Funders | Funder number |
|---|---|
| Defense Advanced Research Projects Agency | MDA972-02-C-0038, MDA972-02-1-0024 |
Keywords
- Broadcast news (BN)
- Conversational telephone speech (CTS)
- Specch-to-text (STT)
ASJC Scopus subject areas
- Acoustics and Ultrasonics
- Electrical and Electronic Engineering
Fingerprint
Dive into the research topics of 'Recent innovations in speech-to-text transcription at SRI-ICSI-UW'. Together they form a unique fingerprint.Cite this
- APA
- Standard
- Harvard
- Vancouver
- Author
- BIBTEX
- RIS