Improving Acoustic Models by Watching Television
Explore this paper's citation graph
Summary
A reliable unsupervised method for identifying accurately transcribed sections of closed-captioned television broadcasts, and it is shown how these segments can be used to train a recognition system.
- Type
- report
- Published
- 1998-03-19
- Cited by
- 17
- References
- 10
- OpenAlex
- https://openalex.org/W1550302919
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:7125952
Keywords
Computer science
References
- Efficient Algorithms for Speech Recognition.
- Improving speech recognition performance via phone-dependent VQ codebooks and adaptive language models in SPHINX-II
- Simultaneous speaker normalisation and utterance labelling using Bayesian/neural net techniques
- The 1994 Abbot hybrid connectionist-HMM large vocabulary recognition system.
- Informedia: news-on-demand multimedia information acquisition and retrieval
- The 1994 HTK large vocabulary speech recognition system
- Cheating with imperfect transcripts
- Language Modeling with Limited Domain Data
- The use of a one-stage dynamic programming algorithm for connected word recognition
- Large vocabulary speech recognition
Cited by
- Segmentation of recordings based on partial transcriptions
- Automated captioning of television programs: development and analysis of a soundtrack corpus
- Reconnaissance automatique de la parole guidée par des transcriptions a priori. (driven decoding for speech recognition system combination)
- Laughter extracted from television closed captions as speech recognizer training data
- New Directions in Video Information Extraction and Summarization
- Environmental audio-visual context awareness
- Transcript synchronization using local dynamic programming
- Combining labeled and unlabeled data with co-training
- Learning to Recognize Speech by Watching Television
- Automated closed-captioning using text alignment
- Multi-Document Summarization and Visualization in the Informedia Digital Video Library
- Generating hypermedia documents from transcriptions of television programs using parallel text alignment
- Informedia - Search and Summarization in the Video Medium
- Using prompts to produce quality corpus for training automatic speech recognition systems
- Audio-to-text alignment for speech recognition with very limited resources
- Bypass Temporal Classification: Weakly Supervised Automatic Speech Recognition with Imperfect Transcripts
- Integrating imperfect transcripts into speech recognition systems for building high-quality corpora
- USING LOCATION INFORMATION FROM SPEECH RECOGNITION OF TELEVISION NEWS BROADCASTS
Related papers
- DETERMINING QUALITY REQUIREMENTS AT THE UNIVERSITIES TO IMPROVE THE QUALITY OF EDUCATION
- Using DataGrid Control to Realize DataBase of Querying in VB6.0
- Study and Two Types of Typical Usage of DataGrid Web Server Control
- PACWON: A parallelizing compiler for workstations on a network
- Bidirectional Sort and Choosing a Row to Update or Delete by Click Any Cell in DataGrid
- Flexible Application of VSFlexGrid
- GMQL: A graphical multimedia query language