Deep Visual-Semantic Alignments for Generating Image Descriptions
Explore this paper's citation graph
- Type
- article
- Published
- 2014-12-06
- Cited by
- 6,113
- References
- 72
- Access
- Open access
- OpenAlex
- https://openalex.org/W2481240925
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:8517067
Keywords
Computer science, Artificial intelligence, Recurrent neural network, Convolutional neural network, Embedding
References
- Midge: Generating Image Descriptions From Computer Vision Detections
- Framing Image Description as a Ranking Task: Data, Models and Evaluation Metrics
- Multimodal learning with deep Boltzmann machines
- Recurrent neural network based language model
- Generating Text with Recurrent Neural Networks
- Unifying Visual-Semantic Embeddings with Multimodal Neural Language Models
- Recurrent Neural Network Regularization
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Composing Simple Image Descriptions using Web-scale N-grams
- Corpus-Guided Sentence Generation of Natural Images
- Microsoft COCO Captions: Data Collection and Evaluation Server
- Show and tell: A neural image caption generator
- A Joint Model of Language and Perception for Grounded Attribute Learning
- From captions to visual concepts and back
- Long-term recurrent convolutional networks for visual recognition and description
- CIDEr: Consensus-based image description evaluation
- BabyTalk: Understanding and Generating Simple Image Descriptions
- I2T: Image Parsing to Text Description
- The Pascal Visual Object Classes (VOC) Challenge
- What Are You Talking About? Text-to-Image Coreference
Cited by
- Towards building a more complex view of the lateral geniculate nucleus: Recent advances in understanding its role
- Learning Two-Branch Neural Networks for Image-Text Matching Tasks
- Interpretable Predictions of Clinical Outcomes with An Attention-based Recurrent Neural Network
- Framing Image Description as a Ranking Task: Data, Models and Evaluation Metrics
- Monocular Object Instance Segmentation and Depth Ordering with CNNs
- Image Captioning with an Intermediate Attributes Layer
- End-to-End People Detection in Crowded Scenes
- Visual Madlibs: Fill in the blank Image Generation and Question Answering
- Deep Knowledge Tracing
- Unsupervised Semantic Parsing of Video Collections
- Multi-Object Classification and Unsupervised Scene Understanding Using Deep Learning Features and Latent Tree Probabilistic Models
- Semantic content-based image retrieval: A comprehensive study
- Convolutional LSTM Network: A Machine Learning Approach for Precipitation Nowcasting
- Are You Talking to a Machine? Dataset and Methods for Multilingual Image Question
- Show, Attend and Tell: Neural Image Caption Generation with Visual Attention
- Hard to Cheat: A Turing Test based on Answering Questions about Images
- Visual Semantic Role Labeling
- Aligning Books and Movies: Towards Story-Like Visual Explanations by Watching Movies and Reading Books
- Invariant backpropagation: how to train a transformation-invariant neural network
- Simple Image Description Generator via a Linear Phrase-based Model
Related papers
- A Study on Performance Improvement of Recurrent Neural Networks Algorithm using Word Group Expansion Technique
- Gated Feedback Recurrent Neural Networks
- Approximating Stacked and Bidirectional Recurrent Architectures with the Delayed Recurrent Neural Network
- Fusion Recurrent Neural Network
- An Exploration of Recurrent Units for Automatic Speech Recognition with RNN based Acoustic Model
- How to Construct Deep Recurrent Neural Networks
- A Structured Self-attentive Sentence Embedding