Neural Discrete Representation Learning
Explore this paper's citation graph
Summary
Pairing these representations with an autoregressive prior, the model can generate high quality images, videos, and speech as well as doing high quality speaker conversion and unsupervised learning of phonemes, providing further evidence of the utility of the learnt representations.
- Type
- preprint
- Published
- 2017-11-02
- Cited by
- 7,851
- References
- 43
- Access
- Open access
- OpenAlex
- https://openalex.org/W2963799213
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:20282961
Keywords
Autoencoder, Computer science, Autoregressive model, Representation (politics), Key (lock)
References
- Deep Boltzmann Machines
- Librispeech: An ASR corpus based on public domain audio books
- Efficient Learning of Domain-invariant Image Representations
- Stochastic Backpropagation and Approximate Inference in Deep Generative Models
- Show and tell: A neural image caption generator
- Auto-Encoding Variational Bayes
- Deep AutoRegressive Networks
- Reinforcement Learning: An Introduction
- Stacked Denoising Autoencoders: Learning Useful Representations in a Deep Network with a Local Denoising Criterion
- Generating Sentences from a Continuous Space
- Estimating or Propagating Gradients Through Stochastic Neurons for Conditional Computation
- Pixel Recurrent Neural Networks
- A Spike and Slab Restricted Boltzmann Machine
- One-shot Learning with Memory-Augmented Neural Networks
- InfoGAN: Interpretable Representation Learning by Information Maximizing Generative Adversarial Nets
- SUPERSEDED - CSTR VCTK Corpus: English Multi-speaker Corpus for CSTR Voice Cloning Toolkit
- Image-to-Image Translation with Conditional Adversarial Networks
- Semi-Supervised Learning with Context-Conditional Generative Adversarial Networks
- Variational Lossy Autoencoder
- PixelVAE: A Latent Variable Model for Natural Images
Cited by
- Fixing a Broken ELBO
- Spatial PixelCNN: Generating Images from Patches
- Hierarchical Text Generation and Planning for Strategic Dialogue
- Distribution Matching in Variational Inference
- On the difficulty of a distributional semantics of spoken language
- Style Tokens: Unsupervised Style Modeling, Control and Transfer in End-to-End Speech Synthesis
- Complex-Valued Restricted Boltzmann Machine for Direct Speech Parameterization from Complex Spectra
- Associative Compression Networks for Representation Learning
- Conditional End-to-End Audio Transforms
- Variational Inference In Pachinko Allocation Machines
- Competitive Training of Mixtures of Independent Deep Generative Models
- Variational Inference for Data-Efficient Model Learning in POMDPs
- A Universal Music Translation Network
- Theory and Experiments on Vector Quantized Autoencoders
- A New Framework for Machine Intelligence: Concepts and Prototype
- Deep Self-Organization: Interpretable Discrete Representation Learning on Time Series
- Gaussian mixture models with Wasserstein distance
- Guided evolutionary strategies: escaping the curse of dimensionality in random search
- Deep Encoder-Decoder Models for Unsupervised Learning of Controllable Speech Synthesis
- Understanding and Improving Interpolation in Autoencoders via an Adversarial Regularizer
Related papers
- TC-VAE: Uncovering Out-of-Distribution Data Generative Factors
- Representation Learning: Recommendation With Knowledge Graph via Triple-Autoencoder
- Representation learning via an integrated autoencoder for unsupervised domain adaptation
- Generative Model for Person Re-Identification: A Review
- Analyzing Big Environmental Audio with Frequency Preserving Autoencoders
- Representation learning: serial-autoencoder for personalized recommendation
- Different latent variables learning in variational autoencoder