Show, Attend and Tell: Neural Image Caption Generation with Visual Attention
Explore this paper's citation graph
Summary
An attention based model that automatically learns to describe the content of images is introduced that can be trained in a deterministic manner using standard backpropagation techniques and stochastically by maximizing a variational lower bound.
- Type
- article
- Published
- 2015-02-10
- Cited by
- 10,911
- References
- 55
- Access
- Open access
- OpenAlex
- https://openalex.org/W1514535095
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:1055111
Keywords
Computer science, Benchmark (surveying), Artificial intelligence, Visualization, Gaze
References
- Midge: Generating Image Descriptions From Computer Vision Detections
- Framing Image Description as a Ranking Task: Data, Models and Evaluation Metrics
- Unifying Visual-Semantic Embeddings with Multimodal Neural Language Models
- Describing Videos by Exploiting Temporal Structure
- Recurrent Neural Network Regularization
- Theano: new features and speed improvements
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Composing Simple Image Descriptions using Web-scale N-grams
- Recurrent Continuous Translation Models
- Generating Sequences With Recurrent Neural Networks
- Stochastic Backpropagation and Approximate Inference in Deep Generative Models
- Corpus-Guided Sentence Generation of Natural Images
- Show and tell: A neural image caption generator
- From captions to visual concepts and back
- Long-term recurrent convolutional networks for visual recognition and description
- Auto-Encoding Variational Bayes
- BabyTalk: Understanding and Generating Simple Image Descriptions
- Control of goal-directed and stimulus-driven attention in the brain
- Long Short-Term Memory
- Dropout: a simple way to prevent neural networks from overfitting
Cited by
- Interactive Sleep Stage Labelling Tool For Diagnosing Sleep Disorder Using Deep Learning
- Image Captioning with an Intermediate Attributes Layer
- Attention-Based Models for Speech Recognition
- Convolutional LSTM Network: A Machine Learning Approach for Precipitation Nowcasting
- Are You Talking to a Machine? Dataset and Methods for Multilingual Image Question
- Video Description Generation Incorporating Spatio-Temporal Features and a Soft-Attention Mechanism
- Visual Semantic Role Labeling
- Aligning Books and Movies: Towards Story-Like Visual Explanations by Watching Movies and Reading Books
- Jointly Modeling Embedding and Translation to Bridge Video and Language
- Describing Videos by Exploiting Temporal Structure
- Modelling serendipity in a computational context
- ReNet: A Recurrent Neural Network Based Alternative to Convolutional Networks
- Exploring Nearest Neighbor Approaches for Image Captioning
- Learning with hidden variables
- Building a Large-scale Multimodal Knowledge Base System for Answering Visual Queries
- Aligning where to see and what to tell: image caption with region-based attention and scene factorization
- End-To-End Memory Networks
- Scalable Bayesian Optimization Using Deep Neural Networks
- Listen, attend and spell: A neural network for large vocabulary conversational speech recognition
- On End-to-End Program Generation from User Intention by Deep Neural Networks
Related papers
- Bottom-Up and Top-Down Attention for Image Captioning and Visual Question Answering
- GloVe: Global Vectors for Word Representation
- Deep Residual Learning for Image Recognition
- ImageNet classification with deep convolutional neural networks
- Learning Phrase Representations using RNN Encoder–Decoder for Statistical Machine Translation
- ROUGE: A Package for Automatic Evaluation of Summaries
- Sequence to Sequence Learning with Neural Networks