Deep neural networks for automatic speech processing: a survey from large corpora to limited data
Explore this paper's citation graph
Summary
This paper investigates state-of-the-art automatic speech recognition systems, as this is the hardest task, and provides an overview of techniques and tasks requiring fewer data, and investigates few-shot techniques by interpreting under-resourced speech as a few- shot problem.
- Type
- article
- Published
- 2020-03-09
- Cited by
- 34
- References
- 63
- Access
- Open access
- OpenAlex
- https://openalex.org/W3009344039
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:212633493
Keywords
Computer science, Speech processing, Task (project management), Speech recognition, Focus (optics)
References
- Fast SVM training based on the choice of effective samples for audio classification
- Librispeech: An ASR corpus based on public domain audio books
- DARPA TIMIT:: acoustic-phonetic continuous speech corpus CD-ROM, NIST speech disc 1-1.1
- Automatic speech recognition for under-resourced languages: A survey
- Severity-Based Adaptation with Limited Data for ASR to Aid Dysarthric Speakers
- Conditional Generative Adversarial Nets
- IEMOCAP: interactive emotional dyadic motion capture database
- SWITCHBOARD: telephone speech corpus for research and development
- Interface Databases: Design and Collection of a Multilingual Emotional Speech Database
- Audio Set: An ontology and human-labeled dataset for audio events
- Prototypical Networks for Few-shot Learning
- VoxCeleb: A Large-Scale Speaker Identification Dataset
- Optimization as a Model for Few-Shot Learning
- Few-Shot Learning with Graph Neural Networks
- Light Gated Recurrent Units for Speech Recognition
- Investigating Generative Adversarial Networks based Speech Dereverberation for Robust Speech Recognition
- Simulating Dysarthric Speech for Training Data Augmentation in Clinical Speech Applications
- VoxCeleb2: Deep Speaker Recognition
- Representation Learning with Contrastive Predictive Coding
- The fifth 'CHiME' Speech Separation and Recognition Challenge: Dataset, task and baselines
Cited by
- Fast offline transformer‐based end‐to‐end automatic speech recognition for real‐world applications
- An exploration of semi-supervised and language-adversarial transfer learning using hybrid acoustic model for hindi speech recognition
- A Two-Level Speaker Identification System via Fusion of Heterogeneous Classifiers and Complementary Feature Cooperation
- Deep Learning Algorithms based Fingerprint Authentication: Systematic Literature Review
- A singular Riemannian geometry approach to Deep Neural Networks I. Theoretical foundations
- Automatic Speech Recognition for Uyghur, Kazakh, and Kyrgyz: An Overview
- A Complete Survey on Generative AI (AIGC): Is ChatGPT from GPT-4 to GPT-5 All You Need?
- A New Global Pooling Method for Deep Neural Networks: Global Average of Top-K Max-Pooling
- An optimized enhanced-multi learner approach towards speaker identification based on single-sound segments
- Bornil: An open-source sign language data crowdsourcing platform for AI enabled dialect-agnostic communication
- A Survey of Audio Classification Using Deep Learning
- QLSTM-based Joint-Training for Noise Robust Hindi Speech Recognition
- Modeling speech processing in case of neurogenic speech and language disorders: neural dysfunctions, brain lesions, and speech behavior
- Lip2Vec: Efficient and Robust Visual Speech Recognition via Latent-to-Latent Visual to Audio Representation Mapping
- Exploring the Role of Convolutional Neural Networks (CNN) in Dental Radiography Segmentation: A Comprehensive Systematic Literature Review
- Digits micro-model for accurate and secure transactions
- Attention Feature Fusion Network via Knowledge Propagation for Automated Respiratory Sound Classification
- An overview of high-resource automatic speech recognition methods and their empirical evaluation in low-resource environments
- Beyond Textual Analysis: Framework for CSAT Score Prediction with Speech and Text Emotion Features
- Advanced Identification of Prosodic Boundaries, Speakers, and Accents Through Multi-Task Audio Pre-Processing and Speech Language Models
Related papers
- Speech signal processing in order to increase recognition of spoken language
- Performance Analysis of Speech Performance Analysis of Speech Recognition using Advanced Algorithm
- Glottal opening instant detection from speech signal
- An investigation into the effect of pitch transformation on children speech recognition
- A microphone array system for speech recognition
- Effect of speech coders on speech recognition performance
- Enhancement and recognition of whispered speech
- A Comparative Study of Feature ExtractionTechniques for Speech Recognition System
- ROBUST RECOGNITION OF SMALL -VOCABULARY TELEPHONE - QUALITY SPEECH