Learning Problem-agnostic Speech Representations from Multiple Self-supervised Tasks
Explore this paper's citation graph
Summary
Experiments show that the proposed improved self-supervised method can learn transferable, robust, and problem-agnostic features that carry on relevant information from the speech signal, such as speaker identity, phonemes, and even higher-level features such as emotional cues.
- Type
- preprint
- Published
- 2019-04-06
- Cited by
- 260
- References
- 46
- Access
- Open access
- OpenAlex
- https://openalex.org/W2935542736
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:102352195
Keywords
Computer science, Encoder, Discriminator, Artificial intelligence, Speech recognition
References
- F0-CONTOURS IN EMOTIONAL SPEECH
- Librispeech: An ASR corpus based on public domain audio books
- The Kaldi Speech Recognition Toolkit
- Delving Deep into Rectifiers: Surpassing Human-Level Performance on ImageNet Classification
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Empirical Evaluation of Gated Recurrent Neural Networks on Sequence Modeling
- A Fast Learning Algorithm for Deep Belief Nets
- Context-Dependent Pre-Trained Deep Neural Networks for Large-Vocabulary Speech Recognition
- Comparison of Parametric Representations for Monosyllabic Word Recognition in Continuously Spoken Se
- Rethinking the Inception Architecture for Computer Vision
- The DIRHA-ENGLISH corpus and related tasks for distant-speech recognition in domestic environments
- Deep Learning of Representations for Unsupervised and Transfer Learning.
- Interface Databases: Design and Collection of a Multilingual Emotional Speech Database
- SUPERSEDED - CSTR VCTK Corpus: English Multi-speaker Corpus for CSTR Voice Cloning Toolkit
- SEGAN: Speech Enhancement Generative Adversarial Network
- Darpa Timit Acoustic-Phonetic Continuous Speech Corpus CD-ROM TIMIT | NIST
- Attentive Convolutional Neural Network Based Speech Emotion Recognition: A Study on the Impact of Input Features, Signal Length, and Acted Speech
- Multi-task Self-Supervised Visual Learning
- Unsupervised Learning of Semantic Audio Representations
- Learning Sight from Sound: Ambient Sound Provides Supervision for Visual Learning
Cited by
- Problem-Agnostic Speech Embeddings for Multi-Speaker Text-to-Speech with SampleRNN
- Generative Pre-Training for Speech with Autoregressive Predictive Coding
- Improving Transformer-based Speech Recognition Using Unsupervised Pre-training
- Learning audio representations via phase prediction
- Speaker-Invariant Affective Representation Learning via Adversarial Training
- Unsupervised learning of cross-modal mappings between speech and text
- From Inference to Generation: End-to-end Fully Self-supervised Generation of Human Face from Speech
- Learning Hierarchical Discrete Linguistic Units from Visually-Grounded Speech
- Visually Guided Self Supervised Learning of Speech Representations
- Multi-Task Self-Supervised Learning for Robust Speech Recognition
- Unsupervised Pre-Training of Bidirectional Speech Encoders via Masked Reconstruction
- Limitations of Weak Labels for Embedding and Tagging
- Towards Learning a Universal Non-Semantic Representation of Speech
- Deep neural networks for automatic speech processing: a survey from large corpora to limited data
- A Comparison of Metric Learning Loss Functions for End-To-End Speaker Verification
- Zero-Shot Multi-Speaker Text-To-Speech with State-Of-The-Art Neural Speaker Embeddings
- Pre-Training Audio Representations With Self-Supervision
- An Early Study on Intelligent Analysis of Speech under COVID-19: Severity, Sleep Quality, Fatigue, and Anxiety
- Vector-Quantized Autoregressive Predictive Coding
- A Further Study of Unsupervised Pretraining for Transformer Based Speech Recognition
Related papers
- Representation Learning with Contrastive Predictive Coding
- BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations
- Mockingjay: Unsupervised Speech Representation Learning with Deep Bidirectional Transformer Encoders
- An Unsupervised Autoregressive Model for Speech Representation Learning
- Librispeech: An ASR corpus based on public domain audio books
- A Simple Framework for Contrastive Learning of Visual Representations