Estimating or Propagating Gradients Through Stochastic Neurons for Conditional Computation
Explore this paper's citation graph
Summary
This work considers a small-scale version of conditional computation, where sparse stochastic units form a distributed representation of gaters that can turn off in combinatorially many ways large chunks of the computation performed in the rest of the neural network.
- Type
- preprint
- Published
- 2013-08-15
- Cited by
- 4,146
- References
- 20
- Access
- Open access
- OpenAlex
- https://openalex.org/W2242818861
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:18406556
Keywords
Differentiable function, Estimator, Computation, Context (archaeology), Computer science
References
- Unsupervised Feature Learning and Deep Learning: A Review and New Perspectives
- Learning representations by back-propagating errors
- Rectified Linear Units Improve Restricted Boltzmann Machines
- Improving neural networks by preventing co-adaptation of feature detectors
- Extracting and composing robust features with denoising autoencoders
- Gradient learning in spiking neural networks by dynamic perturbation of conductances.
- Hierarchical Recurrent Neural Networks for Long-Term Dependencies
- Multivariate stochastic approximation using a simultaneous perturbation gradient approximation
- Deep Sparse Rectifier Neural Networks
- ImageNet classification with deep convolutional neural networks
- Maxout Networks
- Semantic hashing
- Deep Learning of Representations: Looking Forward
- Simple Statistical Gradient-Following Algorithms for Connectionist Reinforcement Learning
- Experimental Gerontology: Acknowledgments
- The Optimal Reward Baseline for Gradient-Based Reinforcement Learning
- Maxout Networks
Cited by
- Deep Directed Generative Autoencoders
- Techniques for Learning Binary Stochastic Feedforward Neural Networks
- Regularizing DNN acoustic models with Gaussian stochastic neurons
- Learning Factored Representations in a Deep Mixture of Experts
- Gradient Estimation Using Stochastic Computation Graphs
- Describing Multimedia Content Using Attention-Based Encoder-Decoder Networks
- Deep AutoRegressive Networks
- Dynamic Capacity Networks
- Behavioral plasticity through the modulation of switch neurons
- Iterative Refinement of Approximate Posterior for Training Directed Belief Networks
- Natural Language Understanding with Distributed Representation
- Conditional Computation in Neural Networks for faster models
- Studies in reinforcement learning and adaptive neural networks
- Compressing Word Embeddings
- Exponentially Increasing the Capacity-to-Computation Ratio for Conditional Computation in Deep Learning
- DoReFa-Net: Training Low Bitwidth Convolutional Neural Networks with Low Bitwidth Gradients
- Minimalist Regression Network with Reinforced Gradients and Weighted Estimates: a Case Study on Parameters Estimation in Automated Welding
- A Deep Learning Framework of Quantized Compressed Sensing for Wireless Neural Recording
- Hierarchical Multiscale Recurrent Neural Networks
- Discrete Variational Autoencoders
Related papers
- Continuous Differentiability in the Context of Generalized Approach to Differentiability
- On binary function continuity,differentiability and may differentiability
- The Complexity Of Nowhere Differentiable Continuous Functions
- On derivatives of fuzzy multi-dimensional mappings and applications under generalized differentiability
- On differentiability of Sobolev functions with respect to the Sobolev norm
- There exist no gaps between Gevrey differentiable and nowhere Gevrey differentiable
- A Necessary and Sufficient Condition on the Differentiability of Binary Functions