Language Modeling with Gated Convolutional Networks
Explore this paper's citation graph
Summary
A finite context approach through stacked convolutions, which can be more efficient since they allow parallelization over sequential tokens, is developed and is the first time a non-recurrent approach is competitive with strong recurrent models on these large scale language tasks.
- Type
- preprint
- Published
- 2016-12-23
- Cited by
- 3,004
- References
- 36
- Access
- Open access
- OpenAlex
- https://openalex.org/W2567070169
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:16119010
Keywords
Computer science, Recurrent neural network, Benchmark (surveying), Sentence, Language model
References
- Hierarchical Probabilistic Neural Network Language Model
- On the importance of initialization and momentum in deep learning
- Recurrent neural network based language model
- Automatic Speech Recognition: A Deep Learning Approach
- Torch7: A Matlab-like Environment for Machine Learning
- Foundations of Statistical Natural Language Processing
- Understanding the difficulty of training deep feedforward neural networks
- One billion word benchmark for measuring progress in statistical language modeling
- Delving Deep into Rectifiers: Surpassing Human-Level Performance on ImageNet Classification
- On the difficulty of training recurrent neural networks
- Improved backing-off for M-gram language modeling
- Long Short-Term Memory
- Three new graphical models for statistical language modelling
- An Empirical Study of Smoothing Techniques for Language Modeling
- Skip-gram Language Modeling Using Sparse Non-negative Matrix Probability Estimation
- Noise-contrastive estimation: A new estimation principle for unnormalized statistical models
- genCNN: A Convolutional Architecture for Word Sequence Prediction
- BlackOut: Speeding up Recurrent Neural Network Language Models With Very Large Vocabularies
- Deep Residual Learning for Image Recognition
- Strategies for Training Large Vocabulary Neural Language Models
Cited by
- Feature Boosting Network For 3D Pose Estimation
- Accelerating Eulerian Fluid Simulation With Convolutional Networks
- Improving Variational Auto-Encoders using Householder Flow
- A Hybrid Convolutional Variational Autoencoder for Text Generation
- Comparative Study of CNN and RNN for Natural Language Processing
- Encoding Sentences with Graph Convolutional Networks for Semantic Role Labeling
- Autoregressive Convolutional Neural Networks for Asynchronous Time Series
- Gated Convolutional Neural Network for Semantic Segmentation in High-Resolution Images
- A Survey of Deep Learning Methods for Relation Extraction
- Supervised Learning of Universal Sentence Representations from Natural Language Inference Data
- Convolutional Sequence to Sequence Learning
- Improving Variational Auto-Encoders using convex combination linear Inverse Autoregressive Flow
- Neural Phrase-based Machine Translation
- Learnable pooling with Context Gating for video classification
- SAM: Semantic Attribute Modulated Language Modeling
- Bottom-Up and Top-Down Attention for Image Captioning and VQA
- Tips and Tricks for Visual Question Answering: Learnings from the 2017 Challenge
- Bottom-Up and Top-Down Attention for Image Captioning and Visual Question Answering
- Sequence-to-Sequence Voice Conversion with Similarity Metric Learned Using Generative Adversarial Networks
- Segmentation-Aware Convolutional Networks Using Local Attention Masks
Related papers
- Study and application of telecom operators system security state baseline
- Analysis of the Baseline Decorrelation and Critical Baseline of Interferometric SAR
- Research of Baseline Implement
- Accounting for baseline targets in NDCs: Issues and options for guidance
- Study on Baseline Evaluation Application in Science and Technology Assessment
- The Baseline: A Social Construction
- Bidirectional recurrent neural network language models for automatic speech recognition
- Language Identification in Short Utterances Using Long Short-Term Memory (LSTM) Recurrent Neural Networks