Semi-Orthogonal Low-Rank Matrix Factorization for Deep Neural Networks
Explore this paper's citation graph
Summary
A factored form of TDNNs (TDNN-F) is introduced which is structurally the same as a TDNN whose layers have been compressed via SVD, but is trained from a random start with one of the two factors of each matrix constrained to be semi-orthogonal.
- Type
- article
- Published
- 2018-09-02
- Cited by
- 523
- References
- 19
- OpenAlex
- https://openalex.org/W2888867175
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:4949673
Keywords
Matrix decomposition, Factorization, Computer science, Artificial neural network, Rank (graph theory)
References
- Training Very Deep Networks
- The Kaldi Speech Recognition Toolkit
- Understanding the difficulty of training deep feedforward neural networks
- Mean-normalized stochastic gradient for large-scale deep learning
- Deep Residual Learning for Image Recognition
- Restructuring of deep neural network acoustic models with singular value decomposition
- On the compression of recurrent neural networks with an application to LVCSR acoustic modeling for embedded speech recognition
- Pronunciation and silence probability modeling for ASR
- A time delay neural network architecture for efficient modeling of long temporal contexts
- Audio augmentation for speech recognition
- Model Compression Applied to Small-Footprint Keyword Spotting
- Purely Sequence-Trained Neural Networks for ASR Based on Lattice-Free MMI
- Low Latency Acoustic Modeling Using Temporal Convolution and LSTMs
- An Exploration of Dropout with LSTMs
- Deep Learning-Based Telephony Speech Recognition in the Wild
- A Pruned Rnnlm Lattice-Rescoring Algorithm for Automatic Speech Recognition
- Neural Network Language Modeling with Letter-Based Features and Importance Sampling
- Highway long short-term memory RNNS for distant speech recognition
Cited by
- Evaluating Code-Switched Malay-English Speech Using Time Delay Neural Networks.
- Training Neural Speech Recognition Systems with Synthetic Speech Augmentation
- The Airbus Air Traffic Control speech recognition 2018 challenge: towards ATC automatic transcription and call sign detection
- The Marchex 2018 English Conversational Telephone Speech Recognition System
- Multiple beamformers with ROVER for the CHiME-5 Challenge
- The NWPU System for CHiME-5 Challenge
- The SHNU system for the CHiME-5 Challenge
- JHU Diarization System Description
- 探討鑑別式訓練聲學模型之類神經網路架構及優化方法的改進 (Discriminative Training of Acoustic Models Leveraging Improved Neural Network Architecture and Optimization Method) [In Chinese]
- Improving LF-MMI Using Unconstrained Supervisions for ASR
- The VOiCES from a Distance Challenge 2019 Evaluation Plan
- Optimisation methods for training deep neural networks in speech recognition
- Acoustic Modeling for Overlapping Speech Recognition: Jhu Chime-5 Challenge System
- Bidirectional LSTM with Extended Input Context
- Lattice-based lightly-supervised acoustic model training
- Semi-supervised acoustic model training for five-lingual code-switched ASR
- Improved low-resource Somali speech recognition by semi-supervised acoustic and language model training
- Lattice-Based Unsupervised Test-Time Adaptation of Neural Network Acoustic Models
- ShrinkML: End-to-End ASR Model Compression Using Reinforcement Learning
- Acoustic Modeling for Automatic Lyrics-to-Audio Alignment
Related papers
- Which scale for parton densities? Highlight from k_t-factorization
- Some rank equalities of some outer inverses of the same matrix
- Spectral Factorization of Rational Matrix Valued Functions
- Generalization of real interval matrices to other fields
- On computing normalized coprime factorization of arbitrary rational matrices
- Existence of a low rank or ℋ︁‐matrix approximant to the solution of a Sylvester equation
- Construction of a full row-rank matrix system for multiple scanning directions in discrete tomography