Skip to main navigation Skip to search Skip to main content

Recent innovations in speech-to-text transcription at SRI-ICSI-UW

  • Andreas Stolcke
  • , Barry Chen
  • , Horacio Franco
  • , Venkata Ramana Rao Gadde
  • , Martin Graciarena
  • , Mei Yuh Hwang
  • , Katrin Kirchhoff
  • , Arindam Mandal
  • , Nelson Morgan
  • , Xin Lei
  • , Tim Ng
  • , Mari Ostendorf
  • , Kemal Sönmez
  • , Anand Venkataraman
  • , Dimitra Vergyri
  • , Wen Wang
  • , Jing Zheng
  • , Qifeng Zhu

Research output: Contribution to journalArticlepeer-review

Abstract

We summarize recent progress in automatic specch-to-text transcription at SRI, ICSI, and the University of Washington. The work encompasses all components of speech modeling found in a state-of-the-art recognition system, from acoustic features, to acoustic modeling and adaptation, to language modeling. In the front end, we experimented with nonstandard features, including various measures of voicing, discriminative phone posterior features estimated by multilayer perceptrons, and a novel phone-level macro-averaging for cepstral normalization. Acoustic modeling was improved with combinations of front ends operating at multiple frame rates, as well as by modifications to the standard methods for discriminative Gaussian estimation. We show that acoustic adaptation can be improved by predicting the optimal regression class complexity for a given speaker. Language modeling innovations include the use of a syntax-motivated almost-parsing language model, as well as principled vocabulary-selection techniques. Finally, we address portability issues, such as the use of imperfect training transcripts, and language-specific adjustments required for recognition of Arabic and Mandarin.

Original languageEnglish (US)
Pages (from-to)1729-1742
Number of pages14
JournalIEEE Transactions on Audio, Speech and Language Processing
Volume14
Issue number5
DOIs
StatePublished - Sep 2006
Externally publishedYes

Funding

Manuscript received October 16, 2005; revised May 30, 2006. This work was supported by the Defense Advanced Research Projects Agency (DARPA) under Contract MDA972-02-C-0038 and Grant MDA972-02-1-0024 (approved for public release, distribution unlimited). The associate editor coordinating the review of this manuscript and approving it for publication was Dr. Alex Acero.

FundersFunder number
Defense Advanced Research Projects AgencyMDA972-02-C-0038, MDA972-02-1-0024

    Keywords

    • Broadcast news (BN)
    • Conversational telephone speech (CTS)
    • Specch-to-text (STT)

    ASJC Scopus subject areas

    • Acoustics and Ultrasonics
    • Electrical and Electronic Engineering

    Fingerprint

    Dive into the research topics of 'Recent innovations in speech-to-text transcription at SRI-ICSI-UW'. Together they form a unique fingerprint.

    Cite this