Speech Waveform Synthesis from MFCC Sequences with Generative Adversarial Networks
Explore this paper's citation graph
- Type
- preprint
- Published
- 2018-04-03
- Cited by
- 55
- References
- 31
- Access
- Open access
- OpenAlex
- https://openalex.org/W2796050583
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:4566442
Keywords
Mel-frequency cepstrum, Computer science, Speech recognition, Cepstrum, Waveform
References
- Introducing CURRENNT: the munich open-source CUDA recurrent neural network toolkit
- Perceptual Properties of Current Speech Recognition Technology
- Theano: new features and speed improvements
- Speaker Recognition by Machines and Humans: A tutorial review
- Linear prediction: A tutorial review
- Quasi Closed Phase Glottal Inverse Filtering Analysis With Weighted Linear Prediction
- Predicting fundamental frequency from mel-frequency cepstral coefficients to enable speech reconstruction.
- Low Bit-Rate Speech Coding Through Quantization of Mel-Frequency Cepstral Coefficients
- Non-parametric techniques for pitch-scale and time-scale modification of speech
- On the inversion of Mel-frequency cepstral coefficients for speech enhancement applications
- An overview of text-independent speaker recognition: From features to supervectors
- Prediction of Fundamental Frequency and Voicing From Mel-Frequency Cepstral Coefficients for Unconstrained Speech Reconstruction
- librosa: Audio and Music Signal Analysis in Python
- Mel-generalized cepstral analysis - a unified approach to speech spectral estimation
- High-pitched excitation generation for glottal vocoding in statistical parametric speech synthesis using a deep neural network
- Least Squares Generative Adversarial Networks
- Using Text and Acoustic Features in Predicting Glottal Excitation Waveforms for Parametric Speech Synthesis with Recurrent Neural Networks
- Non-parallel voice conversion using i-vector PLDA: towards unifying speaker verification and transformation
- Bias and Statistical Significance in Evaluating Speech Synthesis with Mean Opinion Scores
- Sequence-to-Sequence Voice Conversion with Similarity Metric Learned Using Generative Adversarial Networks
Cited by
- Speaker-independent raw waveform model for glottal excitation
- Waveform Generation for Text-to-speech Synthesis Using Pitch-synchronous Multi-scale Generative Adversarial Networks
- Physiological Waveform Imputation of Missing Data using Convolutional Autoencoders
- GELP: GAN-Excited Linear Prediction for Speech Synthesis from Mel-spectrogram
- GlotNet—A Raw Waveform Model for the Glottal Excitation in Statistical Parametric Speech Synthesis
- ASVspoof 2019: Future Horizons in Spoofed and Fake Audio Detection
- Constrained Learned Feature Extraction for Acoustic Scene Classification
- Vocoder-free text-to-speech synthesis incorporating generative adversarial networks using low-/multi-frequency STFT amplitude spectra
- Analysis by Adversarial Synthesis - A Novel Approach for Speech Vocoding
- Deep Residual Neural Networks for Audio Spoofing Detection
- Image-Evoked Affect and its Impact on Eeg-Based Biometrics
- A Speech Reconstruction Algorithm via Iteratively Reweighted ℓ2 Minimization for MFCC Codec
- Neural Source-Filter Waveform Models for Statistical Parametric Speech Synthesis
- Blind Vocoder Speech Reconstruction using Generative Adversarial Networks
- Bridging Mixture Density Networks with Meta-Learning for Automatic Speaker Identification
- The Research of Forensic Voiceprint Identification Based on WMFCC
- Adversarially Training for Audio Classifiers
- A retrieval algorithm for encrypted speech based on convolutional neural network and deep hashing
- Intelligent Instruction-Based IoT Framework for Smart Home Applications using Speech Recognition
- Environment Sound Classification Based on Visual Multi-Feature Fusion and GRU-AWS
Related papers
- Pitch synchronous residual excited speech reconstruction on the MFCC
- Применение частотного маскирования при MFCC-параметризации речи на фоне шумов
- Speech Analysis for Automatic Speech Recognition
- Robust Features for Noisy Speech Recognition using MFCC Computation from Magnitude Spectrum of Higher Order Autocorrelation Coefficients
- Fundamental frequency and voicing prediction from MFCCs for speech reconstruction from unconstrained speech
- Predicting fundamental frequency from mel-frequency cepstral coefficients to enable speech reconstruction.
- Speech Magnitude Spectrum Reconstruction from MFCCs Using Deep Neural Network
- Analysis and prediction of acoustic speech features from mel-frequency cepstral coefficients in distributed speech recognition architectures.