HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units
Explore this paper's citation graph
Summary
The Hidden-Unit BERT (HuBERT) approach for self-supervised speech representation learning, which utilizes an offline clustering step to provide aligned target labels for a BERT-like prediction loss.
- Type
- preprint
- Published
- 2021-06-14
- Cited by
- 5,033
- References
- 64
- Access
- Open access
- OpenAlex
- https://openalex.org/W3169320628
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:235421619
Keywords
Computer science, Cluster analysis, Representation (politics), Consistency (knowledge bases), Artificial intelligence
References
- Librispeech: An ASR corpus based on public domain audio books
- Connectionist Speech Recognition: A Hybrid Approach
- k-means++: the advantages of careful seeding
- A Nonparametric Bayesian Approach to Acoustic Model Discovery
- Unsupervised Training on Large Amounts of Broadcast News Data
- Connectionist temporal classification
- Least squares quantization in PCM
- Applying Convolutional Neural Networks concepts to hybrid NN-HMM model for speech recognition
- Discriminative Training for Large-Vocabulary Speech Recognition Using Minimum Classification Error
- Variational Inference for Acoustic Unit Discovery
- Learning Latent Representations for Speech Generation and Transformation
- Hidden Markov Model Variational Autoencoder for Acoustic Unit Discovery
- Unsupervised Learning of Disentangled and Interpretable Representations from Sequential Data
- Deep Contextualized Word Representations
- Representation Learning with Contrastive Predictive Coding
- Full Bayesian Hidden Markov Model Variational Autoencoder for Acoustic Unit Discovery
- fairseq: A Fast, Extensible Toolkit for Sequence Modeling
- Learning Problem-agnostic Speech Representations from Multiple Self-supervised Tasks
- A Factorial Deep Markov Model for Unsupervised Disentangled Representation Learning from Speech
- Deep Clustering for Unsupervised Learning of Visual Features
Cited by
- Deep neural networks for automatic speech processing: a survey from large corpora to limited data
- Kaizen: Continuously Improving Teacher Using Exponential Moving Average for Semi-Supervised Speech Recognition
- Direct Speech-to-Speech Translation With Discrete Units
- A Longitudinal Normative Dataset and Protocol for Speech and Language Biomarker Research
- SUPERB: Speech processing Universal PERformance Benchmark
- Text-Free Prosody-Aware Generative Spoken Language Modeling
- Deep Multimodal Emotion Recognition on Human Speech: A Review
- Do Infants Really Learn Phonetic Categories?
- Non-Parametric Bayesian Subspace Models for Acoustic Unit Discovery
- Generalization Ability of MOS Prediction Networks
- Distilhubert: Speech Representation Learning by Layer-Wise Distillation of Hidden-Unit Bert
- Improving Pseudo-Label Training For End-To-End Speech Recognition Using Gradient Mask
- Towards High-fidelity Singing Voice Conversion with Acoustic Reference and Contrastive Predictive Coding
- SpeechT5: Unified-Modal Encoder-Decoder Pre-Training for Spoken Language Processing
- Multi-View Self-Attention Based Transformer for Speaker Recognition
- Large-Scale Self-Supervised Speech Representation Learning for Automatic Speaker Verification
- Unispeech-Sat: Universal Speech Representation Learning With Speaker Aware Pre-Training
- Universal Paralinguistic Speech Representations Using self-Supervised Conformers
- Word Order does not Matter for Speech Recognition
- Torchaudio: Building Blocks for Audio and Speech Processing
Related papers
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations
- Librispeech: An ASR corpus based on public domain audio books
- Representation Learning with Contrastive Predictive Coding
- BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
- SpecAugment: A Simple Data Augmentation Method for Automatic Speech Recognition
- Libri-Light: A Benchmark for ASR with Limited or No Supervision
- Mockingjay: Unsupervised Speech Representation Learning with Deep Bidirectional Transformer Encoders