Conformer: Convolution-augmented Transformer for Speech Recognition
Explore this paper's citation graph
Summary
This work proposes the convolution-augmented transformer for speech recognition, named Conformer, which significantly outperforms the previous Transformer and CNN based models achieving state-of-the-art accuracies.
- Type
- preprint
- Published
- 2020-05-16
- Cited by
- 4,274
- References
- 36
- Access
- Open access
- OpenAlex
- https://openalex.org/W3025165719
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:218674528
Keywords
Transformer, Computer science, Convolutional neural network, Language model, Speech recognition
References
- Librispeech: An ASR corpus based on public domain audio books
- Sequence Transduction with Recurrent Neural Networks
- An analysis of noise in recurrent neural networks: convergence and generalization
- Convolutional Neural Networks for Speech Recognition
- Dropout: a simple way to prevent neural networks from overfitting
- Deep convolutional neural networks for LVCSR
- Squeeze-and-Excitation Networks
- Searching for Activation Functions
- State-of-the-Art Speech Recognition with Sequence-to-Sequence Models
- Speech-Transformer: A No-Recurrence Sequence-to-Sequence Model for Speech Recognition
- Convolutional Self-Attention Networks
- Transformer-XL: Attentive Language Models beyond a Fixed-Length Context
- Lingvo: a Modular and Scalable Framework for Sequence-to-Sequence Modeling
- SpecAugment: A Simple Data Augmentation Method for Automatic Speech Recognition
- Attention Augmented Convolutional Networks
- Understanding and Improving Transformer From a Multi-Particle Dynamic System Point of View
- Streaming End-to-end Speech Recognition for Mobile Devices
- Exploring architectures, data and units for streaming end-to-end speech recognition with RNN-transducer
- Learning Deep Transformer Models for Machine Translation
- A Comparative Study on Transformer vs RNN in Speech Applications
Cited by
- Pay Attention to What You Read: Non-recurrent Handwritten Text-Line Recognition
- On the Comparison of Popular End-to-End Models for Large Scale Speech Recognition
- Masked Language Modeling for Proteins via Linearly Scalable Long-Context Transformers
- Continuous Speech Separation with Conformer
- Detecting Acoustic Events Using Convolutional Macaron Net
- Rethinking Attention with Performers
- Swiss Parliaments Corpus, an Automatically Aligned Swiss German Speech to Standard German Text Corpus
- Fairseq S2T: Fast Speech-to-Text Modeling with Fairseq
- Length-Adaptive Transformer: Train Once with Length Drop, Use Anytime with Search
- Universal ASR: Unify and Improve Streaming ASR with Full-context Modeling
- Rethinking Evaluation in ASR: Are Our Models Robust Enough?
- Self-Training and Pre-Training are Complementary for Speech Recognition
- Pushing the Limits of Semi-Supervised Learning for Automatic Speech Recognition
- Developing Real-Time Streaming Transformer Transducer for Speech Recognition on Large-Scale Dataset
- Improving Streaming Automatic Speech Recognition with Non-Streaming Model Distillation on Unsupervised Data
- Transformer-Based End-to-End Speech Recognition with Local Dense Synthesizer Attention
- slimIPL: Language-Model-Free Iterative Pseudo-Labeling
- Perceptual Loss Based Speech Denoising with an Ensemble of Audio Pattern Recognition and Self-Supervised Models
- Improved Mask-CTC for Non-Autoregressive End-to-End ASR
- FastEmit: Low-Latency Streaming ASR with Sequence-Level Emission Regularization
Related papers
- Language Models with Transformers
- Finnish Language Modeling with Deep Transformer Models
- Utilizing Bidirectional Encoder Representations from Transformers for Answer Selection
- Improving Real-time Recognition of Morphologically Rich Speech with Transformer Language Model
- E.T.: Entity-Transformers. Coreference augmented Neural Language Model for richer mention representations via Entity-Transformer blocks